If, Then, Otherwise: Diagnosing Conditional Branching in Vision-Language Navigation

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出CondVLN,一种基于场景图的基准,用于诊断视觉-语言导航中的条件分支问题,通过生成可验证的3D场景图谓词来控制分支执行,以评估代理在感知、逻辑决策等方面的表现。
📝 Abstract
Vision-language navigation agents are often evaluated on their ability to follow route-like instructions toward a fixed goal. Yet, real navigation instructions often depend on observed states of the environment: if a condition holds, then follow one path, otherwise take another. Such instructions require an agent to evaluate scene evidence, select the correct logical branch, and execute the corresponding navigation behavior. Existing evaluations provide limited control over conditional branch execution, making it difficult to determine whether agents fail because of perception, grounding, navigation, or logical decision-making. We introduce CondVLN, a scene-graph-grounded benchmark for diagnosing conditional branching in vision-language navigation. CondVLN programmatically generates instructions whose branch conditions are grounded in verifiable 3D scene-graph predicates, with controlled variation in branch depth, dependency chain length, spatial composition, evidence observability, and instruction horizon. CondVLN contains over 11,500 generated conditional instructions across AI2-THOR, Matterport3D, Gibson, and ReplicaCAD, and evaluates agents using standard VLN metrics and branch-specific diagnostics: Branch Selection Accuracy and Conditional Success Rate. Evaluating four state-of-the-art VLN agents (VLN-Zero, NaVid, NaVILA, and Open-Nav) shows that conditional branching exposes failures that are not captured by standard success rate or path length alone: agents can navigate plausibly while committing to a branch inconsistent with the observed scene condition. We also present a lightweight neurosymbolic branch-selection model that separates condition grounding from navigation execution, improving performance by 2x. CondVLN provides a reusable testbed for measuring whether embodied agents can not only follow instructions, but follow the right instruction under the right condition.
Problem

Research questions and friction points this paper is trying to address.

Conditional Branching
Vision-Language Navigation
Scene Evidence
Logical Decision-Making
Instruction Grounding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conditional Branching
Scene-Grounded Instructions
Branch Selection Accuracy
Neurosymbolic Model
S
Seoyoung Lee
The University of Texas at Austin
Neel P. Bhatt
Neel P. Bhatt
Postdoctoral Fellow, The University of Texas at Austin
RoboticsNeurosymbolic AIComputer VisionState EstimationMotion Forecasting
P
Pranay Samineni
The University of Texas at Austin
C
Cong Liu
Collins Aerospace
S P Sharan
S P Sharan
Ph.D. Student, The University of Texas at Austin
Large Language ModelsMultimodalRoboticsReinforcement LearningNeurosymbolic AI
T
Timothy Barclay
Collins Aerospace
G
Gregory M. Wagner
Collins Aerospace
D
Daniel Milan
The University of Texas at Austin
S
Sandeep Chinchali
The University of Texas at Austin
Ufuk Topcu
Ufuk Topcu
The University of Texas at Austin
autonomycontrolsformal methodslearning
A
Atlas Wang
The University of Texas at Austin