DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses key limitations in aerial vision-and-language navigation—namely, restricted historical context, short planning horizons, and unreliable termination decisions—by introducing a novel approach that integrates a causal memory mechanism, receding-horizon diffusion-based planning, and a lightweight stop-detection module (LiteStop). The method enhances current visual representations with causally aligned historical memory to prevent future information leakage, employs a diffusion policy to predict K-step action sequences while executing only the first step to enable long-horizon planning, and directly estimates stopping probability from action logits. Built upon the Dream-VLA architecture, the proposed system achieves state-of-the-art performance on the OpenFly benchmark, attaining success rates of 32.04% and 29.46% in seen and unseen scenes, respectively, with corresponding SPL scores of 28.22% and 23.54%, and the lowest navigation error among existing methods.
📝 Abstract
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observability. Although recent VLA models offer a promising perception-to-action paradigm, adapting them to aerial navigation remains challenging due to limited historical context, short planning horizons, and unreliable implicit termination. To address these challenges, we propose DreamFly, a diffusion-based aerial VLN framework built on Dream-VLA. DreamFly introduces a causally aligned historical memory that augments the current visual representation using only observations preceding the current decision step, enabling temporal reasoning without future information leakage. We further formulate navigation as receding-horizon diffusion planning, where the policy predicts a $K$-step action chunk but executes only the first action before replanning. This plan-$K$, execute-one strategy uses future actions as auxiliary planning targets while preserving closed-loop visual feedback. Finally, LiteStop estimates the stop probability directly from action logits at the initial all-mask state, decoupling explicit termination from action generation. Experiments on the OpenFly benchmark demonstrate consistent improvements in seen and unseen environments. DreamFly achieves 32.04%/29.46% SR and 28.22%/23.54% SPL on the test-seen/test-unseen splits, respectively, outperforming all compared methods on both metrics while attaining the lowest navigation error. These results demonstrate the effectiveness of jointly modeling historical context, future action structure, and explicit termination for aerial VLN.
Problem

Research questions and friction points this paper is trying to address.

Aerial Vision-Language Navigation
Partial Observability
Temporal Reasoning
Action Planning
Termination Decision
Innovation

Methods, ideas, or system contributions that make the work stand out.

causal memory
receding-horizon diffusion planning
vision-language navigation
explicit termination
aerial navigation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yan Deng
School of Electronic Information Engineering, Xi’an Technological University, Xi’an, China
F
Fei Xu
School of Computer Science and Engineering, Xi’an Technological University, Xi’an, China