Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过将多模态大语言模型直接应用于无人机控制,评估其在指挥、接近、跟踪和搜索任务中的表现,揭示了模型在行动协议上的不足。
📝 Abstract
Multimodal Large Language Models (MLLMs) are strong perceivers of images and video. We ask how far that reach extends into acting: dropping an MLLM directly into a drone's control loop, with its entire action space declared solely in the prompt. Recent systems approach this setting but increasingly narrow the model's decision-making. We widen it back. We introduce DroneCATS-Agent, an architecture where the MLLM is a swappable component, and DroneCATS, a benchmark treating the model as the independent variable. Beyond merely flying toward a pixel, our agent entrusts the model to yaw and search, deliberate when unsure, and self-declare arrival---all without fine-tuning or function-calling schemas. Evaluating frontier and open models across four core capabilities---approaching a visible target, tracking a moving one, searching outside the initial view, and commanding a multi-drone fleet---reveals that even the simplest embodied settings are far from solved. Crucially, to identify what breaks first at the edge, our roster scales down to 2B parameters. The findings expose a stark paradox: it is not the flying that fails. Small open models often navigate into the success radius more reliably than frontier models, yet lose the episode by declaring arrival prematurely or not at all. Multi-drone commanding amplifies this divide, with small models failing by blindly copying a single coordinate across distinct views. Viewed as vision-language-action agents, the models' spatial perception holds up, but their action protocol does not. What separates a deployable edge model from a frontier model is not navigation, but the discipline to sustain a declared protocol and emit the correct terminating action. The open problem is closing this gap at onboard compute costs---yielding a fast model that plans persistently and knows exactly when it is done---and DroneCATS is built to measure that distance.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Large Language Models
Drone Control
Action Protocol
Spatial Perception
Onboard Compute Costs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Large Language Models
Drone Control
Decision-making Expansion
No Fine-tuning
Action Protocol
🔎 Similar Papers
J
Jaewoo Park
Drone AI Team, NAVER Cloud
M
Minyoung Lee
Drone AI Team, NAVER Cloud
S
Sukmin Seo
Drone AI Team, NAVER Cloud
M
Moonbin Yim
Drone AI Team, NAVER Cloud
H
Hyunwook Yoon
Drone AI Team, NAVER Cloud
D
Dohoon Ryu
Drone AI Team, NAVER Cloud
Daehee Kim
Daehee Kim
NAVER Cloud
Deep LearningVision and LanguageOptical Character RecognitionDomain Generalization
M
Myungseo Song
Drone AI Team, NAVER Cloud
J
Jihyuk Byun
Drone AI Team, NAVER Cloud
S
Seunggyu Chang
Drone AI Team, NAVER Cloud
Taeho Kil
Taeho Kil
Drone AI Team, NAVER Cloud
J
Jiseob Kim
Drone AI Team, NAVER Cloud
B
Bado Lee
Drone AI Team, NAVER Cloud
Geewook Kim
Geewook Kim
NAVER Cloud AI & KAIST AI
Large Language ModelsMultimodal LLMsDocument AI