🤖 AI Summary
Goal-directed navigation confronts dual challenges of multimodal sensory fusion and embodied reasoning. This paper proposes a unified analytical framework centered on the “reasoning domain,” systematically surveying nearly 200 recent works to establish, for the first time, a coherent taxonomy of vision-language-audio multimodal navigation methods. It identifies shared computational mechanisms across diverse navigation tasks, distills principles of modality complementarity, and traces paradigmatic evolution in architectural design. Key technical challenges—including cross-modal alignment, joint representation learning, and policy generalization—are rigorously clarified. The study encompasses both simulated and real-world environments, as well as end-to-end and modular architectures, yielding a comprehensive technological landscape. By unifying theoretical insights with scalable implementation pathways, this work lays foundational groundwork for advancing multimodal embodied intelligence.
📝 Abstract
Goal-oriented navigation presents a fundamental challenge for autonomous systems, requiring agents to navigate complex environments to reach designated targets. This survey offers a comprehensive analysis of multimodal navigation approaches through the unifying perspective of inference domains, exploring how agents perceive, reason about, and navigate environments using visual, linguistic, and acoustic information. Our key contributions include organizing navigation methods based on their primary environmental reasoning mechanisms across inference domains; systematically analyzing how shared computational foundations support seemingly disparate approaches across different navigation tasks; identifying recurring patterns and distinctive strengths across various navigation paradigms; and examining the integration challenges and opportunities of multimodal perception to enhance navigation capabilities. In addition, we review approximately 200 relevant articles to provide an in-depth understanding of the current landscape.