ARENA: Automated Red-Teaming for Large Audio Language Models

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that cross-modal safety risks in large audio-language models often evade detection by text-only red teaming. To overcome this, we propose ARENA, a closed-loop automated red teaming framework. This approach introduces a novel audio-grounding mechanism that integrates closed-loop control, MD-Judge reward feedback, and audio variant search to automatically discover joint text-audio attacks eliciting harmful behaviors. Experiments demonstrate that ARENA achieves an attack success rate exceeding 96% across four mainstream models, effectively validating the superiority of its feedback optimization and audio search strategies. Consequently, this work fills a critical gap in the automated evaluation of joint audio-text adversarial attacks, highlighting significant vulnerabilities in current multimodal safety alignment.
📝 Abstract
Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but they also introduce a safety surface that is difficult to expose with text-only red-teaming. We study automated audio-grounded red-teaming, where a text query must remain safe in isolation while the joint text-audio input induces harmful target behavior. We propose ARENA, a closed-loop framework that trains a controller on an independent 2,000case text-audio dataset. MD-Judge supplies training rewards and adaptive search feedback, while a separate, non-adaptive Llama Guard 3 evaluator alone labels final outcomes. On 520 held-out AdvBench objectives, ARENA achieves FDR/PSR of 87.9/100.0%, 71.5/96.3%, 68.1/100.0%, and 75.4/98.5% on Audio Flamingo 3, Qwen2-Audio, MiMo-Audio, and GPTAudio, respectively. Ablations show that feedback-based refinement and audio-variant search substantially improve attack discovery.
Problem

Research questions and friction points this paper is trying to address.

Large Audio Language Models
Automated Red-Teaming
Audio-grounded Attacks
Safety Evaluation
Multimodal Safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automated Red-Teaming
Large Audio Language Models
Closed-loop Framework
Audio-grounded Attack
Adaptive Search Feedback
🔎 Similar Papers
No similar papers found.