The Last Mile of Deepfake Speech Detection: An Industry-Academia Experience Report

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了深度伪造语音检测在实际部署中的问题,通过与商业伙伴合作,提出需要建立共享标准、现实部署基准和易于理解的评分系统。
📝 Abstract
Synthetic speech detection benchmarks now report sub-1% error rates on some in-domain evaluations, yet performance degrades under unseen attacks, channel mismatch, and distribution shift. Based on a three-year effort with Phonexia, a commercial speaker-recognition vendor, we report barriers encountered while building and deploying a detector. Many public benchmarks are not licensed for commercial model development. Real inputs are not four-second clean clips but long, codec-degraded, sometimes partially synthetic recordings. And when a calibrated system returns a log-likelihood ratio of 2.5, no one can tell the customer what it means for their decision. Rather than proposing a new model, we connect these barriers to concrete research and coordination proposals: shared standards for commercially usable datasets, realistic deployment benchmarks, and scores that non-experts can act on. These observations come from one project and should be tested in other settings.
Problem

Research questions and friction points this paper is trying to address.

Deepfake Speech Detection
Performance Degradation
Commercial Deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

commercially usable datasets
realistic deployment benchmarks
scores for non-experts
🔎 Similar Papers