Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii
Mechanistic interpretability (MI) lacks unified, actionable causal evaluation criteria, hindering its advancement. This paper addresses this gap by systematically integrating four classical philosophical accounts of explanation—Bayesian, Kuhnian, Deductive-Nomological, and Mechanistic—into the first multidimensional evaluation framework tailored for mechanistic explanations. We introduce the “compact proof” paradigm: a novel explanatory form that jointly satisfies concision, unification, and generality, and formally specify its generation and verification procedures. Empirical evaluation demonstrates that this paradigm significantly improves explanation quality across diverse models and tasks. Beyond resolving MI’s evaluation bottleneck, our work identifies three foundational research directions: formalizing concision, modeling explanatory unification, and deriving domain-general principles. The framework provides theoretical foundations and methodological tools for building trustworthy AI systems that are monitorable, predictable, and controllable.