Memory Efficient Full-gradient Attacks (MEFA) Framework for Adversarial Defense Evaluations
Existing white-box attacks often resort to approximate gradients when evaluating iterative stochastic purification defenses due to memory constraints, which weakens attack strength and leads to an overestimation of model robustness. This work proposes a memory-efficient full-gradient attack framework that integrates gradient checkpointing with a controllable randomness protocol, enabling—for the first time—exact end-to-end white-box attacks against long-trajectory stochastic defenses such as diffusion- and Langevin-based purification. The method achieves state-of-the-art attack performance under both ℓ∞ and ℓ₂ norms, uncovers vulnerabilities missed by approximate-gradient approaches, and facilitates out-of-distribution robustness analysis, thereby substantially improving the reliability of robustness evaluation.