ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决混合真实性音频深度伪造检测问题,提出ToolDF框架,通过集成工具推理,自适应分析音频场景并提供可解释的判断证据。
📝 Abstract
Audio deepfake detection is commonly formulated as clip-level binary classification of single-domain audio. However, real-world manipulated audio can exhibit mixed authenticity, where genuine and manipulated cues coexist across temporal transitions, overlapping sources, or both. This setting requires not only detecting manipulated audio but also localizing the components that provide evidence for the decision. We propose ToolDF, a tool-integrated reasoning framework for mixed-authenticity audio deepfake detection. ToolDF employs an audio large language model as an orchestrator trained with supervised tool-use trajectories. It adaptively analyzes the audio scene, selectively performs source separation, routes components to domain-specific experts, and aggregates their evidence into an interpretable verdict. We further introduce a mixed-authenticity ADD benchmark covering temporal transitions, acoustic overlaps, and hybrid mixtures. Experimental results show that ToolDF achieves the best overall performance on composite-type detection, achieving macro-F1 gains of 3.72 and 14.39 points over the strongest monolithic baseline and a fixed pipeline, respectively, while providing interpretable evidence localized to temporal regions and acoustic sources. Our source code and dataset are publicly available online.
Problem

Research questions and friction points this paper is trying to address.

audio deepfake detection
mixed authenticity
temporal transitions
acoustic overlaps
Innovation

Methods, ideas, or system contributions that make the work stand out.

tool-integrated reasoning
audio deepfake detection
mixed-authenticity
source separation
domain-specific experts
🔎 Similar Papers
2024-04-22arXiv.orgCitations: 25
💼 Related Jobs
No related jobs found.
T
Taewoo Kim
Multi-Modal Research Center, KETI, South Korea
Y
Young Han Lee
Multi-Modal Research Center, KETI, South Korea
N
Nam In Park
Digital Analysis Section, National Forensic Service, South Korea
C
Chanwoo Kim
Department of Artificial Intelligence, Korea University, South Korea