Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视觉语言模型在实际导航中的部署难题,TAMP-Nav通过2D视觉提示、选择性推理与记忆机制及两级对齐范式来提升导航效率。
📝 Abstract
Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we introduce a Pixel-to-3D Action Formulation (Point) that reformulates navigation into 2D visual prompting. Specifically, the VLM merely selects 2D pixels, which are then projected into 3D coordinates for a low-level SLAM controller. This design naturally aligns embodied execution with the VLM's inherent 2D visual capabilities. Second, we propose an integrated Selective Reasoning and Anchor-Trajectory Memory mechanism (Think and Memorize), which dynamically triggers Chain-of-Thought and retains high-fidelity memory only at critical nodes, compressing redundant trajectories into lightweight Space-Time Indicators, thereby preserving critical historical information and enhancing spatio-temporal perception. Finally, we design an efficient Two-Level Alignment Paradigm (Align) via Group Relative Policy Optimization (GRPO). By superimposing global outcome rewards with fine-grained process rewards, this dense supervision tightly aligns the agent's cognitive planning with physical environmental feedback, endowing the model with adaptive reasoning capabilities. Experiments demonstrate that TAMP-Nav achieves state-of-the-art performance (e.g., 66.2% SR on R2R-CE) with high runtime and sample efficiency (requiring only 90k training trajectories).
Problem

Research questions and friction points this paper is trying to address.

embodied navigation
visual-language models
action space
memory management
reasoning schedule
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pixel-to-3D Action Formulation
Selective Reasoning and Anchor-Trajectory Memory
Two-Level Alignment Paradigm
💼 Related Jobs
No related jobs found.
H
Hongyan Feng
School of Software Technology, Zhejiang University
S
Sunlai Chen
School of Software Technology, Zhejiang University
X
Xuanyu Liu
School of Software Technology, Zhejiang University
Miao Pan
Miao Pan
Professor, Electrical and Computer Engineering, University of Houston
Wireless for AICybersecurity for AIMobile/Edge AI SystemsUnderwater IoT Nets
Y
Yangfan Xie
School of Software Technology, Zhejiang University
Yuxiang Cui
Yuxiang Cui
Zhejiang University
RoboticsReinforcement Learning
Z
Zhongxiang Zhou
Zhejiang Humanoid Robot Innovation Center Co., Ltd.
Rong Xiong
Rong Xiong
Zhejiang University
Robotics
Wenqi Zhang
Wenqi Zhang
Zhejiang University
Language ModelMultimodal LearningEmbodied Agents
Jianwei Yin
Jianwei Yin
Professor of Computer Science and Technology, Zhejiang University
Service ComputingComputer ArchitectureDistributed ComputingAI
Y
Yueting Zhuang
School of Software Technology, Zhejiang University
Xuhong Zhang
Xuhong Zhang
Zhejiang University
LLMVLMVLATrustworthy AI