🤖 AI Summary
Deep learning models in computer vision suffer from opaque decision-making processes and poor interpretability, hindering comprehension by non-experts. To address this, we propose a novel method for generating high-fidelity counterfactual image explanations. Our approach is the first to formally define and optimize a “faithfulness” metric—ensuring explanations strictly reflect the model’s true decision boundary. We integrate gradient-guided adversarial perturbations, latent-space constrained optimization, and a differentiable approximation of the classification boundary within an end-to-end differentiable framework, enabling pixel-level minimal modifications. Evaluated on ImageNet and CUB, our method improves explanation faithfulness by 23.6% over state-of-the-art methods, while significantly enhancing human interpretability. This facilitates fine-grained model diagnosis and debugging.