LG-VLN: A Zero-Shot Vision-and-Language Navigation Framework with LangGraph State Orchestration

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出LG-VLN框架,通过共享视觉特征和基于LangGraph的状态协调解决连续环境下的零样本视觉-语言导航问题,不依赖额外传感器。
📝 Abstract
Continuous-environment vision-and-language navigation (VLN-CE) requires interpreting natural-language instructions in unseen 3D environments and executing continuous low-level actions. Existing methods often depend on LiDAR, panoramic cameras, or extra sensors; separate geometric-mapping and semantic-navigation visual representations can cause long-trajectory spatial-semantic inconsistencies. We propose LG-VLN, a monocular zero-shot framework with shared visual features and LangGraph-based state orchestration. An online feed-forward 3D reconstruction network predicts depth, camera poses, and dense point clouds for agent-pose estimation and global map fusion. Geometry and navigation share dense CleanDIFT features: semantic consistency rejects incorrect inter-frame correspondences, while target-instance constraints define visual references whose similarity combines with local BLIP-2 image-text relevance to form a semantic value map. LangGraph represents instruction parsing, geometric perception, semantic value updates, path planning, action execution, and failure recovery as a directed state graph with conditional transitions, persistent state, and modular recovery mechanisms. On a fixed 550-episode subset of the R2R-CE val-unseen split, LG-VLN achieves 21.3% success and 12.1% success weighted by path length. Ablations show shared semantic features improve navigation, further boosted by combining visual similarity and image-text relevance. Results establish shared visual representations and explicit state orchestration as effective for zero-shot VLN-CE using monocular RGB alone. Code will be publicly released for reproducibility.
Problem

Research questions and friction points this paper is trying to address.

vision-and-language navigation
continuous environment
spatial-semantic inconsistency
natural-language instructions
Innovation

Methods, ideas, or system contributions that make the work stand out.

zero-shot
shared visual features
LangGraph
semantic consistency
monocular RGB
🔎 Similar Papers
No similar papers found.
J
Jianhe Zhao
School of Geodesy and Geomatics, Wuhan University, Wuhan, China
Y
Yanhua Qiu
School of Geodesy and Geomatics, Wuhan University, Wuhan, China
Zhiyu Zhang
Zhiyu Zhang
Postdoc, Carnegie Mellon University
Machine LearningOptimizationStatistics
Zibo Zhao
Zibo Zhao
Hunyuan, Tencent; ShanghaiTech
J
Jinhua Xie
School of Geodesy and Geomatics, Wuhan University, Wuhan, China