AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization

๐Ÿ“… 2026-08-02
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenges of high latency and out-of-distribution (OOD) collisions faced by lightweight drones navigating complex 3D environments. The authors propose an efficient and robust vision-language navigation approach that leverages high-fidelity visual inputs to drive a compact 2B-parameter vision-language model. A key innovation is a zero-cost automated Direct Preference Optimization (DPO) framework, which exploits state rollback and privileged intervention in physics-based simulation to automatically generate collision-avoidance preference dataโ€”eliminating the need for manual annotation. Experimental results demonstrate a 49.16% success rate in previously unmapped scenes, significantly reducing collision rates and setting a new state-of-the-art in autonomous drone navigation. The study further reveals that perceptual fidelity plays a more critical role in navigation performance than linguistic reasoning capability.
๐Ÿ“ Abstract
Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist end-to-end paradigms show great promise but typically rely on massive language models containing billions of parameters, incurring prohibitive latency for real-world edge deployment. In this paper, we challenge this parameter-heavy reliance. Comprehensive cross-scale evaluations reveal the critical insight that perception quality fundamentally outweighs language reasoning capacity. We demonstrate that a lightweight 2B model equipped with high-fidelity visual inputs completely matches the overall success rates of massive 7B baselines. However, this minimalist policy exposes a fundamental robustness flaw inherent to pure Behavior Cloning (BC). Lacking explicit negative feedback, the agent fails to internalize robust spatial constraints and exhibits alarming collision rates in out-of-distribution (OOD) scenarios. To overcome this vulnerability without relying on unscalable human annotations, we propose AeroDPO, a zero-cost automated Direct Preference Optimization pipeline driven by deterministic physical simulation state rollback. Upon detecting collisions, the system autonomously rewinds the environment to extract causal reasoning errors as rejected actions, applies decoupled privileged interventions to synthesize collision-avoidance preferred maneuvers, and leverages an offline vision language inspector to filter visual ambiguities. By equipping our 2B model with this automated data flywheel, AeroDPO boosts success rates to 49.16% on unmapped scenarios while drastically suppressing collision rates, establishing a new SOTA for autonomous aerial agents.
Problem

Research questions and friction points this paper is trying to address.

UAV-VLN
lightweight navigation
robustness
collision avoidance
out-of-distribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

lightweight UAV navigation
high-fidelity perception
automated preference optimization
Direct Preference Optimization (DPO)
vision-language navigation
๐Ÿ”Ž Similar Papers
Peng Xu
Peng Xu
Huazhong University of Science and Technology
Data SecurityCryptography
C
Chengcheng Wang
University of Electronic Science and Technology of China, Chengdu, China; Shenzhen Institute for Advanced Study, UESTC, Shenzhen, China
S
Shaohua Wan
University of Electronic Science and Technology of China, Chengdu, China; Shenzhen Institute for Advanced Study, UESTC, Shenzhen, China