Time-Aware Assistive Navigation

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
Can interactive vision-and-language agents learn not just what to say but also \textbf{\textit{when}} to say it? Current language models rarely plan over whether and when to realize a real-time response to a user. However, providing accurate and timely support for human decision-making, such as when guiding visually impaired individuals through urban environments, requires careful real-time responsiveness--poorly timed responses can distract users or add unnecessary cognitive load. As a machine intelligence challenge for Multimodal Large Language Model (MLLM)-based agents, we introduce a large-scale multimodal benchmark for an egocentric, assistive navigation task in complex outdoor environments. Using this benchmark, we uncover a fundamental limitation of off-the-shelf MLLMs in delivering safe and time-sensitive navigation instructions, even with model fine-tuning on substantial amounts of data. We then demonstrate that a simple yet effective modification of the model, including direct supervision to predict the underlying reason for each instruction, yields significant performance gains across open-loop, closed-loop, and sim-to-real generalization settings. However, our analysis highlights persistent challenges in temporal reasoning, safety-critical object awareness, and relational and distance understanding. To advance the development of scalable assistive agents, we will release our simulation, benchmark, and code (available at the project website: https://timeli-icra.github.io/).
Problem

Research questions and friction points this paper is trying to address.

Time-Aware
Assistive Navigation
Multimodal Large Language Model
Real-time Responsiveness
Safety-critical
Innovation

Methods, ideas, or system contributions that make the work stand out.

time-aware navigation
multimodal benchmark
temporal reasoning
safety-critical awareness
model modification
🔎 Similar Papers
No similar papers found.