LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决VLN中数据需求大和计算存储开销高的问题,提出LookStep框架,通过语言中心的未来状态建模和事件驱动的滚动记忆来提高导航效率。
📝 Abstract
Vision-Language Navigation (VLN) requires an embodied agent to follow natural-language instructions in unseen environments. Recent progress has been largely driven by Multimodal Large Language Models (MLLMs). Existing methods follow a next-step action prediction paradigm, supervising only the expert action, which requires a high quantity of data for training. They also rely on cognitive maps, accumulated historical frames, or external 3D tools to maintain states, leading to high computational and memory overhead. To realize resource efficiency VLN, we propose LookStep, a unified end-to-end framework that combines Language Centric Future State Modeling and Event Driven Rolling Memory that uses language labels to generate coarse-grained navigation progress and future states for each candidate action, while autonomously deciding whether to write each observation into a bounded rolling memory with a semantic role. We validate LookStep empirically. On VLN-CE tasks, LookStep outperforms existing methods under the same training settings, achieving a 49.7\% success rate on R2R-CE Val-Unseen with better memory efficiency and less data usage. Code and model is available at https://github.com/kunyang-YU/LookStep.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Navigation
Efficiency
Resource Efficiency
Data Usage
Memory Overhead
Innovation

Methods, ideas, or system contributions that make the work stand out.

Language Centric Future State Modeling
Event Driven Rolling Memory
Vision-Language Navigation
Efficient Resource Utilization
Kun-Yang Yu
Kun-Yang Yu
LAMDA Group, Nanjing University
Machine Learning
Yingzhe Li
Yingzhe Li
Samsung Research America
Wireless CommunicationsStochastic Geometry5GLTEWi-Fi
Hongyu Xu
Hongyu Xu
Research Scientist, Meta Reality Labs
Spatial PerceptionGenAIMultimodalRoomPlan
Shi-Yu Tian
Shi-Yu Tian
Nanjing University
machine learning
Z
Zhi Zhou
National Key Laboratory for Novel Software Technology, Nanjing University; School of Artificial Intelligence, Nanjing University
Y
Yang Chen
National Key Laboratory for Novel Software Technology, Nanjing University; School of Intelligence Science and Technology, Nanjing University
Ming Yang
Ming Yang
Nanjing University
S
Sheng Wang
TermiTech
Q
Qing Yu
TermiTech
Lan-Zhe Guo
Lan-Zhe Guo
LAMDA Group, Nanjing University
Machine Learning
Yu-Feng Li
Yu-Feng Li
Professor, Nanjing University
Machine Learning