Seeing What Matters: Visual Cue Guided Video Planning for Generalizable Robot Navigation

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过结合视觉线索引导的视频规划和特定实体的逆动力学模型,解决了机器人长距离导航和精确动作转换的问题。
📝 Abstract
Generative video models can serve as a promising backbone for robot navigation by predicting future observations as video plans. Recent approaches often condition video planning on short-horizon guidance and recover geometric waypoints through scene reconstruction, leaving longer-horizon planning and precise video-to-action translation less explored. We present CueNav, a video model-based navigation framework combining visual cue guided video planning with an embodiment-specific Inverse-Dynamics Model (IDM). As visual cues, we use a Bird's-Eye View (BEV) map to convey global task context and retain part of the robot body in the egocentric observation to expose embodiment context. These cues guide the video planner, while the IDM translates dense flow fields extracted from the video plan into robot actions. With the visual cue encoding global task context, CueNav achieves nearly 2x higher success in maze navigation than planning without the cue. The body-aware view with the IDM enables precise navigation with 70% success in a narrow passage where comparison methods largely fail to complete the task. We further demonstrate zero-shot semantic-conditioned navigation and deployment of the same video planner across different robot platforms. Our results show that visual cue-guided video planning with embodiment-specific action grounding paves the way toward a generalizable navigation framework for longer-horizon planning and embodiment-aware control. Additional results and code are available on our project website: https://cuenav.github.io.
Problem

Research questions and friction points this paper is trying to address.

video planning
robot navigation
longer-horizon planning
visual cue
embodiment-aware control
Innovation

Methods, ideas, or system contributions that make the work stand out.

CueNav
Visual Cue Guided Video Planning
Inverse-Dynamics Model (IDM)
Bird's-Eye View (BEV)
🔎 Similar Papers
No similar papers found.
H
Hojin Lee
Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich, 80797 Munich, Germany
S
Sizhe Lester Li
Computer Science and Artificial Intelligence Laboratory (CSAIL), Massachusetts Institute of Technology, Cambridge, MA 02139, USA
Maximilian Hilger
Maximilian Hilger
Research Associate, Technical University of Munich
RoboticsRadarLocalizationMappingSLAM
S
Susie Lu
Computer Science and Artificial Intelligence Laboratory (CSAIL), Massachusetts Institute of Technology, Cambridge, MA 02139, USA
Achim J. Lilienthal
Achim J. Lilienthal
Full Professor & MIRMI Deputy Director at TU Munich / Guest Professor at Örebro University
Robot PerceptionMobile RoboticsArtificial IntelligenceMobile Robot OlfactionMathematics
V
Vincent Sitzmann
Computer Science and Artificial Intelligence Laboratory (CSAIL), Massachusetts Institute of Technology, Cambridge, MA 02139, USA
D
Daniel A. Duecker
Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich, 80797 Munich, Germany