Unifying Formal Explanations: A Complexity-Theoretic Perspective
This work establishes a unified theoretical framework for analyzing the computational complexity of sufficient and contrastive explanations in machine learning models. It introduces a general probabilistic value function whose minimization subsumes both explanation types, enabling rigorous analysis through combinatorial optimization and computational complexity theory. The key contribution lies in demonstrating, for the first time, that under global explanation settings, this value function exhibits monotonicity, submodularity, or supermodularity—properties that guarantee efficient polynomial-time computability for a broad class of explanations. In stark contrast, even highly simplified variants become NP-hard in local explanation settings. These results provide a unified theoretical foundation and clear computational feasibility criteria for explainability across diverse model classes, including neural networks and decision trees.