Hybrid Audio Detection Using Fine-Tuned Audio Spectrogram Transformers: A Dataset-Driven Evaluation of Mixed AI-Human Speech

📅 2025-05-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Emerging AI-based voice attacks increasingly involve mixed audio—comprising authentic speech, fully synthetic speech, voice clones, and hybrid combinations—posing novel security threats to voice authentication systems. Method: To address this, we introduce the first comprehensive benchmark dataset covering all four audio categories and propose a novel hybrid audio detection framework based on fine-tuned Audio Spectrogram Transformer (AST). Unlike conventional binary classification approaches, our method pioneers a hybrid acoustic pattern modeling paradigm, integrating spectrogram-based representation learning, transfer learning, and controllable hybrid audio synthesis with precise annotation. Contribution/Results: Experiments demonstrate that our approach achieves 97% accuracy on hybrid audio detection—significantly outperforming existing baselines. This work fills critical gaps in both data resources and modeling methodologies for hybrid voice detection, establishing a reproducible benchmark and an effective technical pathway to enhance the robustness of speaker verification systems against sophisticated voice spoofing attacks.

Technology Category

Application Category

📝 Abstract
The rapid advancement of artificial intelligence (AI) has enabled sophisticated audio generation and voice cloning technologies, posing significant security risks for applications reliant on voice authentication. While existing datasets and models primarily focus on distinguishing between human and fully synthetic speech, real-world attacks often involve audio that combines both genuine and cloned segments. To address this gap, we construct a novel hybrid audio dataset incorporating human, AI-generated, cloned, and mixed audio samples. We further propose fine-tuned Audio Spectrogram Transformer (AST)-based models tailored for detecting these complex acoustic patterns. Extensive experiments demonstrate that our approach significantly outperforms existing baselines in mixed-audio detection, achieving 97% classification accuracy. Our findings highlight the importance of hybrid datasets and tailored models in advancing the robustness of speech-based authentication systems.
Problem

Research questions and friction points this paper is trying to address.

Detecting hybrid audio mixing human and AI-generated speech segments
Addressing security risks in voice authentication from advanced cloning
Improving detection accuracy for complex mixed acoustic patterns
Innovation

Methods, ideas, or system contributions that make the work stand out.

Constructed hybrid audio dataset with diverse samples
Fine-tuned Audio Spectrogram Transformer models
Achieved 97% accuracy in mixed-audio detection
🔎 Similar Papers
K
Kunyang Huang
Kean University
B
Bin Hu
Kean University