AT-ADD: A Benchmark and Challenge for Robust and All-Type Audio Deepfake Detection

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出AT-ADD基准测试,旨在解决全类型音频深度伪造检测问题,通过两大赛道评估模型在不同条件下的鲁棒性及泛化能力。
📝 Abstract
Recent audio generation models can synthesize high-fidelity speech, environmental sound, singing voice, and music, creating new risks for multimedia trust. Existing audio deepfake detection (ADD) benchmarks remain predominantly speech-centric and often underrepresent realistic channel variation and diverse audio types. This paper presents AT-ADD, a large-scale benchmark and challenge designed to evaluate both robust speech deepfake detection and all-type audio deepfake detection. Track 1 evaluates binary speech detection under unseen generators, diverse recording conditions, signal perturbations, and replay effects. Track 2 evaluates type-agnostic real/fake detection over speech, sound, singing, and music when the audio type is unknown at test time. We detail the dataset construction, evaluation protocol, and reproducible baselines, and analyze the final systems submitted to the ACM Multimedia 2026 Grand Challenge. The strongest official baseline obtains 76.73% and 79.47% Macro-F1 on the Track 1 and Track 2 evaluation sets, respectively, whereas the winning challenge systems reach 90.71% and 96.10%. Beyond aggregate rankings, sample-level analysis of the top five submissions examines generator- and type-level difficulty, cross-system error complementarity, and ranking stability. The results show that large-scale self-supervised representations, condition-aware augmentation, multi-crop inference, and structured fusion or routing are central to generalization, while generator-specific robustness and consistent performance across diverse audio types remain unresolved.
Problem

Research questions and friction points this paper is trying to address.

Audio Deepfake Detection
Robust Speech
Diverse Audio Types
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio deepfake detection
all-type audio
self-supervised representation
condition-aware augmentation
multi-crop inference
🔎 Similar Papers
2024-04-22arXiv.orgCitations: 25
💼 Related Jobs
No related jobs found.
Yuankun Xie
Yuankun Xie
PhD Candidate, Communication University of China
Audio Deepfake DetectionDomain GeneralizationOut-of-Distribution DetectionNeural Audio Codec
H
Haonan Cheng
Communication University of China, Beijing, China
J
Jiayi Zhou
Machine Intelligence, Ant Group, Shanghai, China
X
Xiaoxuan Guo
Communication University of China, Beijing, China; also with Ant Group
T
Tao Wang
Machine Intelligence, Ant Group, Shanghai, China
C
Changhao Zhang
Machine Intelligence, Ant Group, Shanghai, China
J
Jian Liu
Machine Intelligence, Ant Group, Shanghai, China
W
Weiqiang Wang
Machine Intelligence, Ant Group, Shanghai, China
Ruibo Fu
Ruibo Fu
Associate Professor,CASIA
AIGCLMMIntelligent speech interactionDeepfake detection
Xiaopeng Wang
Xiaopeng Wang
Institute of Automation, Chinese Academy of Sciences
Fake Audio DetectionText To SpeechSpeech Large Model
Hengyan Huang
Hengyan Huang
Pursuing a degree in Intelligent Science and Technology at the Communication University of China.
AIGCMLLMAudio-Visual Processing
X
Xiaoying Huang
Communication University of China, Beijing, China
Long Ye
Long Ye
Communication University of China
Multimedia Signal ProcessingArtificial Intelligence
Guangtao Zhai
Guangtao Zhai
Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI EvaluationDisplays