One MLLM, One Call: Efficient Zero-Shot Vision-and-Language Navigation via Spatial-Aware Waypoints

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出O2C-Nav框架,通过单次调用大型模型和使用结构化航点生成器解决零样本视觉-语言导航中的高延迟和计算开销问题。
📝 Abstract
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to navigate unseen environments by following natural language instructions. Current zero-shot VLN-CE methods either rely on pre-trained waypoint predictors or require multiple queries to large models per step. To address prohibitive inference latency and computational overhead, we propose O2C-Nav, an efficient zero-shot navigation framework that calls only a single large model once per decision step. Our approach introduces a training-free structured waypoint generator and a novel abstract representation that projects sparse, history-aware candidate waypoints directly onto RGB images as visual markers. The MLLM selects a waypoint or generates a fallback target bounding box at each step, while a low-level Fast Marching Method (FMM) planner converts the selected target into an executable collision-free path. This paradigm provides the model with concrete spatial perception and explicit memory while significantly reducing the visual processing load. Extensive evaluations on the R2R-CE and RxR-CE benchmarks demonstrate that O2C-Nav outperforms current state-of-the-art zero-shot methods, highlighting its great potential for real-time robotic deployment. Code is available at https://github.com/kkpsq/O2C-Nav-Code.
Problem

Research questions and friction points this paper is trying to address.

Vision-and-Language Navigation
Zero-Shot
Inference Latency
Computational Overhead
Innovation

Methods, ideas, or system contributions that make the work stand out.

zero-shot navigation
single large model call
structured waypoint generator
abstract representation
Fast Marching Method (FMM)
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Shiqi Pan
College of Electronics and Information Engineering, Shenzhen University, Shenzhen 518060, China
Qi Zheng
Qi Zheng
Shenzhen University
artificial intelligencemachine learning
H
Hanqin Sun
College of Electronics and Information Engineering, Shenzhen University, Shenzhen 518060, China
Youjian Zhang
Youjian Zhang
the University of Sydney
computer visionimage processing
Daquan Feng
Daquan Feng
Shenzhen University
D2DEdge IntelligenceSpatial ComputingGreen Communications...
X
Xu Wang
College of Computer Science and Software Engineering, Shenzhen University, Shenzhen 518060, China