MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing audio deepfake detection methods fail to account for the fundamental differences between speech and environmental sounds in terms of generation mechanisms and artifact characteristics, leading to biased evaluations and poor generalization in multi-source, independently manipulated scenarios. To address this limitation, this work proposes MADBench, the first component-aware benchmark for audio deepfake detection, which explicitly treats speech and environmental sounds as distinct acoustic components. MADBench introduces a multi-scenario test set encompassing independent manipulations of both components and provides a unified evaluation protocol for state-of-the-art detectors, general-purpose audio encoders, and multimodal large language models. Experimental results reveal that environmental sound forgeries are more easily detectable than speech forgeries, that current detectors consistently underperform on both components, and that manipulated environmental sounds significantly degrade the detection performance for forged speech.
📝 Abstract
Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet existing research either focuses on visual manipulation, addresses speech detection in isolation, or conflates speech and non-speech audio as a single undifferentiated audio stream, overlooking the distinct forensic challenges posed by background audio. This conflation is consequential: the two acoustic components arise from fundamentally different generative mechanisms, exhibit distinct artifact profiles, and pose different challenges to detection systems. We introduce MADBench, the first benchmark that treats speech and environmental audio as distinct acoustic components, enabling component-aware evaluation of audio deepfake detection across independently manipulated forgery sources. We benchmark representative state-of-the-art detectors and multimodal large language models under a unified protocol. Our experiments reveal that environmental audio manipulation is more detectable than synthetic speech across general-purpose encoders, while existing pretrained detectors fail on both acoustic components, and manipulated environmental audio asymmetrically degrades speech deepfake detection, findings entirely invisible under the single-label paradigm of prior benchmarks. MADBench establishes a rigorous foundation for future research into robust, component-aware audio deepfake detection.
Problem

Research questions and friction points this paper is trying to address.

audio deepfake detection
modality-aware
speech manipulation
environmental audio
forensic challenges
Innovation

Methods, ideas, or system contributions that make the work stand out.

modality-aware
audio deepfake detection
component-aware benchmark
environmental audio manipulation
MADBench
🔎 Similar Papers
2024-04-22arXiv.orgCitations: 25