From Interpretability Methods to Interpretable Models

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文主张将可解释AI研究重点从方法转向模型,通过现有工具评估模型的可理解性及计算内容,以增强对模型的信任。
📝 Abstract
More than a decade in, explainable AI (XAI) for computer vision has assembled a mature toolbox: attribution, feature visualization, concept-based, and circuit-based methods. Yet almost all of the field's effort has gone into building and comparing these methods, and little into the question they were meant to answer---how interpretable are our models, and are we making progress as they evolve? We argue for shifting the field's focus from methods to models, along two complementary lines. One is already within reach: existing tools let us characterize and compare what different models represent and compute. The other is harder, and largely neglected: whether a model can actually be understood by the humans who rely on it---the independent evaluators on whom trust and certification depend, not the experts confirming what they already expect. It can only be measured, not inferred. We review why the toolbox is mature enough to support both, survey the thin body of work comparing models, draw a parallel to systems neuroscience, and close with a model-centric XAI agenda.
Problem

Research questions and friction points this paper is trying to address.

explainable AI
interpretability
models
human understanding
trust
Innovation

Methods, ideas, or system contributions that make the work stand out.

model-centric
interpretability
XAI
human-understandable