Audio Deepfake Detection Using Temporal Coherence Analysis

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于CLAP嵌入的时间一致性分析框架,通过计算音频片段嵌入的余弦相似性并提取统计特征来训练轻量级集成分类器,以区分真实与合成音频。
📝 Abstract
The proliferation of AI-generated audio (so-called "deepfake" audio) poses significant threats to information integrity, from voice cloning fraud to synthetic music copyright disputes. We present a temporal coherence analysis framework built upon Contrastive Language-Audio Pretraining (CLAP) embeddings that spans speech, instrumental music, and music with vocals. By computing pairwise cosine similarities between audio segment embeddings and extracting statistical features from the resulting distributions, we train lightweight ensemble classifiers that reliably distinguish authentic from synthetic audio. Our work provides an interpretable, computationally efficient alternative to common deep learning methods while still achieving competitive performance across speech and music domains. Further, we reveal two notable empirical findings about audio deepfakes: (1) a feature-label inversion phenomenon in which 21 of 29 statistical features reverse their discriminative direction between training and in-the-wild deployment, and (2) a speech--music direction reversal in which entropy discriminates in opposite directions for speech and music deepfakes.
Problem

Research questions and friction points this paper is trying to address.

audio deepfake
information integrity
voice cloning fraud
synthetic music copyright disputes
Innovation

Methods, ideas, or system contributions that make the work stand out.

Temporal Coherence Analysis
CLAP Embeddings
Lightweight Ensemble Classifiers
Feature-Label Inversion
🔎 Similar Papers
2024-04-22arXiv.orgCitations: 25