MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

📅 2026-07-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing multimodal drone benchmarks inadequately assess decision compliance and safety under perceptual degradation and linguistic ambiguity, as they overlook the coupling among physical evidence, safety protocols, and action risk. This work proposes the first offline, protocol-constrained vision–language–action (VLA) decision-level evaluation benchmark, establishing a structured diagnostic framework encompassing 12 dimensions and 17 task categories to jointly evaluate protocol adherence, action safety, and multimodal evidence arbitration. Through semantic scoring and modality ablation studies across 17 models, the best protocol-to-decision score achieved is only 0.5141, with an average dimensional accuracy as low as 0.1599. These findings reveal the critical influence of visual and textual inputs on action selection and expose fundamental deficiencies in current models regarding constraint extraction and cross-modal trust calibration.
📝 Abstract
Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degraded observations and ambiguous language. Existing UAV and multimodal benchmarks evaluate perception, navigation, collaboration, and reasoning, but few assess whether physical evidence, protocol constraints, and action risk remain coupled during critical decisions. We introduce MulRobBench, an offline, protocol-conditioned benchmark for Vision-Language-Action (VLA) UAV agents in smart-city environments. MulRobBench integrates real UAV multimodal observations, protocol-level security policies, and action-level cyber-physical safety into a unified evaluation framework. The benchmark contains 3,024 samples spanning 17 task taxonomy nodes and 12 scoring dimensions across four stages: operational context understanding, multimodal evidence arbitration, degradation-aware reasoning, and risk-aware action planning. Evaluation combines semantic scoring with structural diagnostics, including policy compliance, format compliance, unsafe actions, parsing failures, and dimension-level validity. Across 17 multimodal models, the best semantic protocol-decision score reaches only 0.5141, while the best strict mean scoring-dimension accuracy is 0.1599. A controlled 20-anchor modality-ablation study changes 4-15 action selections per model, confirming that both visual and textual inputs influence decisions. Analysis identifies modality-trust selection, constraint extraction, glare, missing data, and operator shorthand as the primary causes of decision instability. MulRobBench provides a reproducible benchmark for trustworthy multimodal UAV decision making under realistic operational constraints.
Problem

Research questions and friction points this paper is trying to address.

multimodal UAV agents
decision-level benchmark
security-policy compliance
cyber-physical safety
protocol-conditioned evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

multimodal UAV agents
decision-level benchmark
security-policy compliance
risk-aware action planning
protocol-conditioned evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Belal S. Alsinglawi
Belal S. Alsinglawi
Zayed University, Swinburne University of Technology
Internet of ThingsAI AlgorithmsSecurity and Privacy
Weizheng Wang
Weizheng Wang
Hong Kong Polytechnic University
Information SecurityApplied CryptographyBlockchain
J
Junyi Wu
School of Computer Science and Engineering, University of Emergency Management, Beijing 101601, China
Yi Jiang
Yi Jiang
Department of Civil and Environmental Engineering, Hong Kong Polytechnic University
Environmental NanotechWater TreatmentEnvironmental ChemistryAerosol Technology
L
Lianhai Lin
School of Computer Science and Engineering, University of Emergency Management, Beijing 101601, China
Merouane Debbah
Merouane Debbah
KU 6G Center, Khalifa University, Centralesupelec
6GLarge Language ModelsAIRandom Matrix TheoryGame Theory
I
Izzat Alsmadi
Department of Computational, Engineering and Mathematical Sciences, College of Arts and Sciences, Texas A&M University–San Antonio, San Antonio, TX 78224 USA