Fake It Until You Break It: On the Adversarial Robustness of AI-generated Image Detectors
This study addresses the insufficient adversarial robustness of AI-generated image detectors in real-world settings, where they are vulnerable to black-box attacks and common social media degradations (e.g., JPEG compression, resizing, color distortion), enabling malicious misuse for disinformation and undermining democratic trust. We conduct the first systematic evaluation of mainstream detectors under combined black-box adversarial perturbations and realistic degradations, revealing that state-of-the-art models suffer over 40% accuracy degradation without model access. To mitigate this, we propose a lightweight CLIP-enhanced defense grounded in zero-shot detection and black-box transfer attack modeling—requiring no retraining or fine-tuning. Our method preserves original detection performance while reducing adversarial success rates by 76%, substantially restoring practical utility and robustness on real platforms. This work delivers a deployable, trustworthy solution for AI-generated content authentication.