Language Orthogonalization for Zero-Shot Cross-Lingual Audio Deepfake Detection

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过语言正交化方法解决跨语言音频深度伪造检测中的零样本迁移问题,减少了自监督语音模型中的语言依赖性结构干扰。
📝 Abstract
Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-supervised speech models (S3Ms), these backbones encode language-dependent structure that confounds spoof cues. We address this confound through language orthogonalization, a target-free ridge map that removes S3M variation projected onto continuous language-identification (LID) embeddings. Across six languages, six S3M backbones, and all Leave-N-Out settings, it consistently reduces EER across unseen languages. Cross-lingual EER correlates with LID-space distance, where orthogonalization yields larger gains for more distant transfers.
Problem

Research questions and friction points this paper is trying to address.

Audio Deepfake Detection
Cross-Lingual
Language Orthogonalization
Self-Supervised Speech Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Language Orthogonalization
Self-supervised Speech Models
Cross-lingual Audio Deepfake Detection
Language-identification Embeddings
🔎 Similar Papers
No similar papers found.