Towards counterfactual and contrastive explainability and transparency of DCNN image classifiers
To address the insufficient decision transparency of deep convolutional neural networks (DCNNs) in image classification, this paper proposes the first unified interpretability framework that jointly models counterfactual perturbations and class-contrastive mechanisms, generating human-understandable “why-this-not-that” explanations. Methodologically, it integrates gradient-guided counterfactual search, contrastive attention mask optimization, a differentiable semantic editing module, and a class-aware loss function—ensuring both local faithfulness and global consistency while enabling fine-grained attribution and controllable semantic editing. Experiments on ImageNet and CUB-200 demonstrate significant improvements: +23.6% in explanation fidelity and +31.2% in user trustworthiness. The generated explanations exhibit strong semantic plausibility and visual verifiability.