AdaGeoVLN: Selective Geometry Across Representation Depth and Navigation Time for Vision-Language Navigation

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出AdaGeoVLN方法,通过层次化融合和导航感知记忆解决视觉-语言导航中几何特征使用及历史几何证据保留问题。
📝 Abstract
Vision-language navigation requires aligning language with visual observations while maintaining spatial understanding over time. Geometry foundation models (GFMs) expose intermediate representations throughout their hierarchy, but how navigation policies should use these features and retain historical geometric evidence remains unresolved. We introduce \method{}, a streaming VLN framework that addresses these questions across \textbf{representation depth} and \textbf{navigation time}. Hierarchical GFM--VLM fusion couples earlier, intermediate, and later GFM representations to successive policy stages instead of repeatedly injecting a terminal feature. Navigation-aware GFM memory retains historical VGGT global-attention KV states according to instruction relevance, geometric confidence, and transition novelty under a bounded per-layer budget. Retained states provide geometric context for subsequent observations before fusion with the policy. Across R2R-CE and RxR-CE, \method{} achieves strong performance using a single RGB stream without additional navigation-specific external data. Controlled ablations show that multi-depth coupling substantially outperforms repeated terminal-feature injection at matched fusion locations. Bounded navigation-aware retention preserves navigation performance while considerably reducing GFM-KV memory relative to larger-memory temporal retention. These findings support jointly examining the geometric representations exposed to the policy and the historical evidence retained for future inference. Code will be released upon acceptance at https://humanoid-research.github.io/adageovln/.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Navigation
Geometry Foundation Models
Historical Geometric Evidence
Representation Depth
Navigation Time
Innovation

Methods, ideas, or system contributions that make the work stand out.

Selective Geometry
Representation Depth
Navigation Time
Hierarchical GFM--VLM Fusion
Navigation-aware GFM Memory
Quan-Dung Pham
Quan-Dung Pham
VinMotion, Inc., Vietnam
Anh Dao
Anh Dao
Undergraduate Student, Michigan State University
Vision-languageMultimodal LLMEmbodied AILLM
D
Danh Vinh Le
VinMotion, Inc., Vietnam
N
Nguyen Viet Tri Pham
VinMotion, Inc., Vietnam
T
The-Anh Nguyen
VinMotion, Inc., Vietnam
Zhirui Dai
Zhirui Dai
UC San Diego
Robotics
Y
Yiyu Chen
VinMotion, Inc., Vietnam
T
Tuyen P. Le
VinMotion, Inc., Vietnam
T
Truong Nguyen
VinMotion, Inc., Vietnam
Quan Nguyen
Quan Nguyen
University of Southern California
ControlRoboticsOptimization