A Robot System for Indoor Environment Question Answering With Cognitive Map Leveraging Vision-Language Models
Journal
IEEE Transactions on Automation Science and Engineering
Journal Volume
23
Start Page
6090
End Page
6101
ISSN
1545-5955
1558-3783
Date Issued
2026-03-09
Author(s)
Abstract
The paper introduces 'Environment Question Answering (EnvQA),' an advanced task derived from Embodied Question Answering (EQA), aiming to enhance the practicality of robotic systems in real-world scenarios. Unlike EQA, which involves repeated exploration of familiar environments for answering questions, EnvQA integrates spatial memory and user feedback to address these limitations. The EnvQA system is designed to autonomously navigate, answer queries, and store environmental information using a cognition-inspired cognitive map with intermediate features of Vision-language models (VLMs). Additionally, the paper presents a new zero-shot image-image matching method, ViTPR, which utilizes large-scale vision transformers to achieve state-of-the-art performance in Visual Place Recognition (VPR) on the highly challenging Nordland dataset. Experimental results also demonstrate the effectiveness of global localization and question-answering in both simulated and physical environments. In summary, we successfully use VLMs to enhance the robot's ability to understand and associate human questions with the environment. Note to Practitioners - This work aims to develop an autonomous mobile robot system that allows people to use natural language to have the robot assist in confirming the status of various people, objects, and events in different home environments, and to respond in natural language as well. Our approach utilizes autonomously constructed cognitive maps to directly answer questions, eliminating the need for repeated exploration for each user query. The robot will navigate to the relevant location and provide responses based on current observations only when the user requests re-verification. We believe this system can serve as a convenient platform to provide effective assistance in long-term care, hospital rounds, and even general home caregiving environments, demonstrating significant potential in healthcare. It is important to note that current limitations include the system's dependency on the quality of VLMs and the need for more robust real-time processing capabilities. Future research should focus on addressing these limitations and exploring scalable applications across diverse environments.
Subjects
cognitive map
cognitive robot
embodied question answering
Environment question answering
visual place recognition
Publisher
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal article
