Unmasking Deepfakes: Leveraging Augmentations and Features Variability for Deepfake Speech Detection

📅 2025-01-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address the poor robustness of deepfake speech detection in low-resource, low-data-language scenarios, this paper proposes an end-to-end hybrid architecture integrating a self-supervised feature extractor with a lightweight classification head. We introduce two novel techniques: multi-level spectrogram masking augmentation and a compressed sensing–inspired pretraining mechanism—jointly enabling audio-level and feature-level regularization. The method effectively mitigates monolingual, few-shot constraints and demonstrates strong generalization across unseen codecs and previously unencountered spoofing attacks. On the ASVSpoof2024 closed-set Track 1 benchmark, our approach achieves an EER of 4.37%; substituting the pretrained feature extractor further reduces the EER to 3.39%, establishing a new state-of-the-art performance.

Technology Category

Application Category

📝 Abstract
The detection of deepfake speech has become increasingly challenging with the rapid evolution of deepfake technologies. In this paper, we propose a hybrid architecture for deepfake speech detection, combining a self-supervised learning framework for feature extraction with a classifier head to form an end-to-end model. Our approach incorporates both audio-level and feature-level augmentation techniques. Specifically, we introduce and analyze various masking strategies for augmenting raw audio spectrograms and for enhancing feature representations during training. We incorporate compression augmentations during the pretraining phase of the feature extractor to address the limitations of small, single-language datasets. We evaluate the model on the ASVSpoof5 (ASVSpoof 2024) challenge, achieving state-of-the-art results in Track 1 under closed conditions with an Equal Error Rate of 4.37%. By employing different pretrained feature extractors, the model achieves an enhanced EER of 3.39%. Our model demonstrates robust performance against unseen deepfake attacks and exhibits strong generalization across different codecs.
Problem

Research questions and friction points this paper is trying to address.

Deepfake Detection
Minority Languages
Limited Data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Deep Learning System
Audio Feature Extraction
ASVSpoof5 Performance
🔎 Similar Papers
Inbal Rimon
Inbal Rimon
Ben-Gurion University
O
Oren Gal
University of Haifa, Haifa, Israel
H
H. Permuter
Ben Gurion University, Be’er Sheva, Israel