GroundingVLN: Reasoning and Acting with Grounding for Vision-Language Navigation

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决视觉-语言导航中语义推理与空间执行的连接问题,提出GroundingVLN方法,通过视觉锚定作为推理和行动的共享接口。
📝 Abstract
Although vision-language models (VLMs) possess strong visual understanding and reasoning capabilities, existing vision-and-language navigation (VLN) agents struggle to connect semantic reasoning with spatial execution. Two coupled gaps remain in this connection, as intermediate reasoning is not explicitly anchored to visual evidence and high-level decisions lack precise spatial goals to guide low-level motion. Cognitive science suggests that human navigation bridges these levels hierarchically by anchoring cognition to relevant landmarks and guiding locomotion toward spatial goals. Motivated by this principle, we propose GroundingVLN, which uses visual grounding as a shared interface between reasoning and action. GroundingVLN first reasons with grounding by anchoring task-relevant visual evidence to precise image locations throughout structured reasoning. It then acts through grounding by predicting a progress-aligned pixel goal that a geometric planner translates into primitive actions. To learn these capabilities, we construct GroundingCOTVLN-188K, a dataset of temporally aligned grounded reasoning traces, and introduce Grounded and Execution-Aware Reinforcement Learning (GEAR), which aligns grounded reasoning and spatial decisions with downstream execution. Experiments demonstrate that GroundingVLN achieves state-of-the-art performance (69.9% SR on R2R-CE and 75.1% SR on RxR-CE) with high sample efficiency, using just 0.9% as much training data as the strongest baseline. It also generalizes strongly across datasets, attaining 59.9% SR on RxR-CE when trained solely on R2R, a gain of 20.1% over the strongest baseline.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Navigation
Semantic Reasoning
Spatial Execution
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual grounding
structured reasoning
progress-aligned pixel goal
GEAR (Grounded and Execution-Aware Reinforcement Learning)
💼 Related Jobs
No related jobs found.
K
Kailing Li
School of Computer Science and Technology, East China Normal University
Yu Han
Yu Han
Professor of Chemistry, South China University of Technology
NanomaterialsElectron MicroscopyCatalysis
Tianwen Qian
Tianwen Qian
East China Normal University
MultimediaVision and LanguageEmbodied AI
Y
Yuqian Fu
King Abdullah University of Science and Technology
Jingyu Gong
Jingyu Gong
Shanghai Jiao Tong University
3D Computer Vision
J
Jiangming Shi
School of Computer Science and Technology, East China Normal University
X
Xiaoling Wang
School of Computer Science and Technology, East China Normal University