AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations in drone vision-and-language navigation stemming from the absence of explicit scene modeling and future spatial reasoning capabilities. The authors propose a current-to-future spatial map imagination framework that learns a structured representation of the current environment from multi-view observations, jointly supervised by current map reconstruction and future trajectory prediction. A structured causal attention mechanism is introduced to propagate present spatial knowledge into future map inference, enabling collaborative prediction of the next 3D waypoint. Furthermore, a novel cross-spatial planning consistency loss is designed to align the predicted map trajectory with expert action directions. The method achieves state-of-the-art performance on the OpenUAV and AerialVLN-S benchmarks, and ablation studies confirm the effectiveness of individual components and the overall robustness of the framework.
📝 Abstract
Unmanned Aerial Vehicle Vision-Language Navigation (UAV-VLN) requires agents to follow language instructions, infer spatial structure from sparse multi-view observations, and execute feasible 3D motion in complex outdoor environments. Despite recent progress with large language models, most existing methods still map vision-language inputs directly to actions, providing limited explicit scene grounding and future-aware spatial reasoning. We propose AirForesight, a current-to-future spatial map imagination framework for UAV-VLN. AirForesight first learns a structured current-map representation from multi-view observations. This representation is jointly supervised by current-map reconstruction and future-trajectory prediction, encouraging it to encode both present scene structure and future motion intent. Under structured causal attention, the current spatial knowledge is propagated to future-map reasoning, and the resulting current and future representations are aggregated to predict the next 3D waypoint. To make spatial imagination more relevant to navigation, we introduce a cross-space planning consistency loss that encourages directional agreement between the predicted map-space trajectory and the expert action direction derived from the ground-truth waypoint displacement. Experiments on OpenUAV and AerialVLN-S, together with extensive ablations, demonstrate strong performance and support the effectiveness and stability of the proposed framework.
Problem

Research questions and friction points this paper is trying to address.

UAV-VLN
spatial reasoning
scene grounding
future-aware navigation
3D motion planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

spatial map imagination
cross-space planning consistency
structured causal attention
future-aware reasoning
UAV-VLN
🔎 Similar Papers
No similar papers found.