Disentangled Global-Local Feature Learning with E-Branchformer for Audio Deepfake Detection

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为应对语音合成技术对认证系统的威胁,提出基于E-Branchformer的新架构,通过并行分支捕捉全局和局部特征,并结合深度卷积和挤压激励模块增强鉴别能力,以实现音频深伪检测。
📝 Abstract
The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervised speech representations for audio deepfake detection. Our model employs parallel branches to simultaneously capture global contextual dependencies through multi-head self-attention and local temporal patterns through convolutional processing. To enhance discriminative capability, we integrate depthwise convolution and Squeeze-and-Excitation modules that enrich the classification token with refined patch token information after feature merging. Extensive experiments on ASVspoof 2021 LA, DF, and In-the-Wild datasets demonstrate state-of-the-art performance with equal error rates of 0.88%, 1.85%, and 6.30% respectively, substantially outperforming existing methods. Comprehensive ablation studies validate that the dual-branch architecture provides complementary discriminative information, Squeeze-and-Excitation Aggregation significantly improves SSL feature integration, and the combination of DWConv and SE modules is critical for effective class token enhancement. The superior performance on real-world scenarios demonstrates strong generalization capability to diverse acoustic conditions and unseen spoofing attacks.
Problem

Research questions and friction points this paper is trying to address.

audio deepfake detection
voice synthesis technologies
speech-based authentication systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

E-Branchformer
self-supervised speech representations
multi-head self-attention
depthwise convolution
Squeeze-and-Excitation
🔎 Similar Papers
P
Phuong Tuan Dat
Department of Electrical and Computer Engineering, National University of Singapore
H
Ho Bao Thu
School of Communication and Information Technology, Hanoi University of Science and Technology
N
Nguyen Tran Trung
School of Communication and Information Technology, Hanoi University of Science and Technology
P
Pham Viet Hoang
School of Communication and Information Technology, Hanoi University of Science and Technology
Nguyen Thi Thu Trang
Nguyen Thi Thu Trang
Lecturer & Researcher, School of Information and Communication Technology, Hanoi University of
Speech SynthesisSpeaker RecognitionSpeech TechnologyNatural Language Processing