Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignnment Score for Retinal Disease Classification
Current deep learning models for retinal disease classification suffer from limited interpretability and an inability to reliably link predictions to clinically relevant lesion regions, hindering their clinical deployment. To address this, this work proposes CounterFundus, a novel framework that leverages CycleGAN to generate healthy counterfactual images corresponding to pathological fundus images. Lesion localization is achieved through difference maps between original and counterfactual images. The study introduces CCAS—a unified evaluation metric combining Spearman correlation, Intersection over Union (IoU), and pointing accuracy—to quantitatively assess the spatial alignment between counterfactual explanations and classifier saliency maps. Experiments demonstrate that the generated counterfactuals exhibit strong alignment with classification evidence across all CCAS dimensions, and that CCAS-guided counterfactual data augmentation significantly enhances downstream classification performance.