CRAF: Cross-View Residual-Aware Fusion for Deepfake Speech Detection

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决深度伪造语音检测中未见过的攻击问题,提出CRAF方法,通过跨视图残差感知融合增强自监督学习模型与听觉大语言模型的信息整合,提高检测鲁棒性。
📝 Abstract
Recent advances in speech synthesis and voice conversion have made deepfake speech increasingly realistic, making generalization to unseen spoofing attacks a critical challenge. Pretrained speech and audio models offer a promising direction for improving robustness to such unseen attacks. Self-supervised learning (SSL) models capture fine-grained, low-level acoustic characteristics, whereas Auditory Large Language Models (ALLMs) provide higher-level contextual representations. These complementary views can provide useful cues for improving generalization to unseen attacks. However, direct fusion does not explicitly disentangle information shared across the two views from view-specific complementary information, limiting effective cross-view integration. To address this, we propose CRAF, a cross-view residual-aware fusion framework that uses ALLM-guided cross-view attention to enrich SSL representations and adopts ALLM as a high-level reference to separate ALLM-explainable information from complementary SSL residual information. The residual is selectively refined through adaptive gating and integrated through SSL-primary fusion. Experiments on ASVspoof 5 show that CRAF with Kimi-Audio achieves an EER of 5.96% and a minDCF of 0.1192, demonstrating robustness to unseen spoofing attacks.
Problem

Research questions and friction points this paper is trying to address.

Deepfake Speech
Unseen Spoofing Attacks
Cross-View Fusion
Self-Supervised Learning
Auditory Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-View Residual-Aware Fusion
Self-supervised Learning (SSL)
Auditory Large Language Models (ALLMs)
Adaptive Gating
Deepfake Speech Detection
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
M
Minh-Xuan Phan
Japan Advanced Institute of Science and Technology
K
Khalid Zaman
Japan Advanced Institute of Science and Technology
C
Candy Olivia Mawalim
Japan Advanced Institute of Science and Technology
Masashi Unoki
Masashi Unoki
JAIST
Auditory modelspeech signal processing