AT-ADD: All-Type Audio Deepfake Detection Challenge Summary

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of universal audio deepfake detection and robustness in real-world transmission channels by establishing the first comprehensive benchmark spanning speech, singing, music, and environmental sounds. We propose a novel detection paradigm integrating large-scale self-supervised representations, multi-crop inference, and structured routing fusion to effectively overcome generalization bottlenecks across heterogeneous audio types. Experimental results demonstrate that this combined strategy achieves Macro-F1 scores of 90.71% and 96.10% in Track 1 and Track 2, respectively, validating its efficacy in complex acoustic scenarios. Furthermore, this work systematically summarizes design patterns for cross-modal universal detection, providing critical insights and a foundational reference for advancing all-type audio forgery identification research.
📝 Abstract
This paper summarizes the ACM Multimedia 2026 AT-ADD Grand Challenge on all-type audio deepfake detection. AT-ADD contains two tracks: robust speech deepfake detection under realistic acoustic and channel variations, and type-agnostic detection over speech, environmental sound, singing voice, and music. We describe the challenge tasks, dataset and evaluation-set design, official leaderboard results, and common design patterns observed in participating systems. The best Track 1 system achieved 90.71% Macro-F1 on the final evaluation set, while the best Track 2 system achieved 96.10% Macro-F1. The final submissions show that strong systems commonly combine large-scale self-supervised audio representations, data augmentation, multi-crop inference, and structured fusion or routing. The results also reveal remaining challenges in generalization to unseen generators, robustness to realistic speech-domain distortions, and balanced performance across heterogeneous audio types.
Problem

Research questions and friction points this paper is trying to address.

Audio Deepfake Detection
Type-agnostic Detection
Robustness
Generalization
Heterogeneous Audio
Innovation

Methods, ideas, or system contributions that make the work stand out.

All-Type Audio Deepfake Detection
Self-Supervised Audio Representations
Type-Agnostic Detection
Structured Fusion
Robustness
🔎 Similar Papers
2024-04-22arXiv.orgCitations: 25