α3-Bench: A Unified Benchmark of Safety, Robustness, and Efficiency for LLM-Based UAV Agents over 6G Networks
Current evaluation frameworks struggle to comprehensively assess the safety, protocol compliance, and task effectiveness of large language model (LLM)-driven drone agents under the dynamic constraints of 6G networks. To address this gap, this work proposes α³-Bench, a novel benchmark that integrates safety, robustness, and efficiency into a unified α³ evaluation metric. Built upon a multi-turn conversational control framework, α³-Bench features a dual-action-layer architecture enabling tool invocation and multi-agent collaboration, while incorporating 6G network emulation—including latency, jitter, and packet loss—and tool consistency verification. Evaluation across 17 prominent LLMs on a dataset of 113k dialogues reveals that although most models achieve high task success rates under nominal conditions, their robustness and communication efficiency degrade significantly under impaired 6G network conditions.