VTInstructor: Visual Trajectory Prompting for Navigation Instruction Generation in Continuous Environments

πŸ“… 2026-08-15
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of generating navigation instructions from continuous RGB video streams where trajectory cues are absent. We propose the first map-free instruction generation framework for continuous-environment Vision-and-Language Navigation (VLN). By employing visual trajectory prompting to externalize implicit geometry, combined with keyframe extraction, signal injection, and VT-GRPO training calibration, our method eliminates reliance on scene reconstruction. Experiments demonstrate that this framework establishes a new state-of-the-art on the R2R-CE benchmark with significant CIDEr improvements. Furthermore, it achieves a downstream navigation success rate of 63.3%, representing a 14.7 percentage point increase. These results confirm the framework’s effectiveness in overcoming instruction generation difficulties in map-free environments, offering a robust solution for continuous VLN without explicit spatial maps.
πŸ“ Abstract
Navigation instruction generation from ego-centric RGB video in continuous environments is an important yet challenging task for human-robot interaction and scalable dataset construction. Prior instruction generators assume discrete viewpoint graphs with panoramic observations, where trajectory structure is explicit; in continuous environments, however, the agent receives only a dense RGB stream, making trajectory cues difficult to recover. We propose VTInstructor, the first VLN instruction generation framework for continuous environments. Our key idea is to convert implicit trajectory geometry into explicit visual trajectory prompts: EDTC condenses long RGB trajectories into navigation-critical keyframes, VTP overlays path, turn, and goal cues onto these anchors, VTMod injects the resulting trajectory signals into the visual encoder, and VT-GRPO further calibrates this spatial injection during training, all without requiring a navigation graph, pre-built map, or scene reconstruction. On the challenging R2R-CE and RxR-CE Val Unseen benchmarks, VTInstructor sets a new state of the art across all standard NLG metrics, surpassing the strongest baseline by +0.357 CIDEr and +0.109 CIDEr, respectively. Beyond automatic metrics, VTInstructor-generated instructions raise a frozen follower's success rate to 63.3%, a +14.7 percentage-point gain over the best competing instruction source, and provide consistent data augmentation gains of +3 SR points on downstream navigation tasks.
Problem

Research questions and friction points this paper is trying to address.

Navigation Instruction Generation
Continuous Environments
Ego-centric RGB Video
Visual Language Navigation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Trajectory Prompting
Continuous Environment VLN
VT-GRPO
Navigation Instruction Generation
Keyframe Condensation
πŸ”Ž Similar Papers
No similar papers found.
Haolin Yang
Haolin Yang
University of Chicago
large language modelsnatural language processing
Yuxing Long
Yuxing Long
Peking University
Embodied Intelligence
Z
Zihan Yang
CFCS, School of Computer Science, Peking University
H
Hao Dong
CFCS, School of Computer Science, Peking University