Interpretable All-Type Audio Deepfake Detection with Audio LLMs via Frequency-Time Reinforcement Learning
This work addresses the challenge of generalizing audio deepfake detection across diverse audio types—including speech, environmental sounds, singing, and music—where existing methods struggle to balance performance and interpretability. The authors propose a two-stage training framework based on Audio Large Language Models (ALLMs). First, they construct interpretable supervision signals via a frequency-time structured chain-of-thought (CoT) with automatic annotation. Subsequently, they perform reinforcement fine-tuning by integrating supervised fine-tuning (SFT) with a novel Frequency-Time Grouped Relative Policy Optimization (FT-GRPO). The resulting model achieves state-of-the-art performance across all audio forgery detection tasks while generating human-interpretable reasoning grounded in frequency-time features, effectively mitigating reward hacking and hallucination issues.