ICON Decomposition: Multivariate Concept-Level Explanations of Deep Representations for Model Auditing

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出ICON分解方法,通过考虑所有概念及结果间的相互作用来准确评估深度模型中各概念的重要性,解决模型审计中的误导性关联问题。
📝 Abstract
Deep neural networks often exploit spurious associations in their training data, a failure known as shortcut learning. Concept-based explainability methods screen for shortcuts by testing whether concepts such as a patient's sex or scanner settings can be decoded from a network layer. Because each concept is evaluated in isolation, these methods can mistake correlations between concepts as evidence that the model uses them. We introduce ICON decomposition, which instead quantifies how much of a layer's variance each concept explains after accounting for all other concepts and the outcome. On synthetic data with known ground truth, ICON recovers concept importance more accurately than seven alternative baseline methods. On skin-lesion and brain-imaging models, it isolates the concepts on which a model genuinely relies, quantifies the representation unexplained by any of the supplied concepts, and yields sparse explanations that we validate by retraining and out-of-distribution testing.
Problem

Research questions and friction points this paper is trying to address.

deep neural networks
shortcut learning
concept-based explainability
spurious associations
Innovation

Methods, ideas, or system contributions that make the work stand out.

ICON Decomposition
Concept Importance
Model Auditing
Deep Representations
🔎 Similar Papers
No similar papers found.
R
Roshan Prakash Rane
Hertie Institute for AI in Brain Health, University of Tübingen, Tübingen, Germany.
M
Marco Simnacher
Chair of Statistics, Humboldt-Universität zu Berlin, Berlin, Germany.
M
Manuel Pfeuffer
Chair of Statistics, Humboldt-Universität zu Berlin, Berlin, Germany.
M
Marc-Andre Schulz
Hertie Institute for AI in Brain Health, University of Tübingen, Tübingen, Germany.
N
Nys Tjade Siegel
Hertie Institute for AI in Brain Health, University of Tübingen, Tübingen, Germany.
Maximilian Dreyer
Maximilian Dreyer
Explainable AI Group, Fraunhofer Heinrich Hertz Institute
Explainable AI (XAI)InterpretabilityArtificial IntelligenceComputer Vision
Frederik Pahde
Frederik Pahde
Fraunhofer Heinrich Hertz Institute
Machine LearningExplainable AIComputer VisionFew-shot LearningMultimodality
Wojciech Samek
Wojciech Samek
Professor at TU Berlin, Head of AI Department at Fraunhofer HHI, BIFOLD Fellow
Deep LearningInterpretabilityExplainable AITrustworthy AIFederated Learning
Sonja Greven
Sonja Greven
Chair of Statistics, Humboldt-Universität zu Berlin
functional data analysislongitudinal data analysisjoint modelsflexible regression modelsbiostatistics
K
Kerstin Ritter
Hertie Institute for AI in Brain Health, University of Tübingen, Tübingen, Germany.