FoeGlass: Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors
This work addresses the limitations of current audio deepfake detection models, whose evaluation relies on manually curated data and thus fails to efficiently uncover critical blind spots. To overcome this, the authors propose FoeGlass—the first black-box, automated red-teaming framework tailored for text-to-speech systems. FoeGlass leverages the in-context learning capability of large language models to explore the input space and generate adversarial audio samples capable of evading detection. By incorporating diversity-guided prompt design to mitigate mode collapse, FoeGlass substantially enhances attack efficacy in black-box settings, reducing detector false negative rates by up to 94% while demonstrating strong cross-model transferability. Furthermore, fine-tuning detectors on FoeGlass-generated data improves their robustness by as much as 41%.