Class-Aware Reinforcement Learning for Counterfactual Explanation Generation
This work proposes a class-aware reinforcement learning framework to enhance the efficiency and quality of counterfactual explanation generation. By incorporating the model’s predicted class information into the state representation—a novel design in this domain—the proposed approach guides the policy to more effectively explore counterfactual instances that satisfy validity, sparsity, and proximity constraints. The integration of class awareness significantly accelerates policy convergence and improves reward optimization. Empirical evaluation across seven diverse datasets demonstrates that the method generates a greater number of high-quality explanations with fewer training episodes compared to existing approaches. Furthermore, feature importance analyses using SHAP and LIME confirm that class information plays a critical role in guiding action selection during the generation process.