Medical AI Encodes a"Feeling of Error": Verifying Cancer Segmentation via Internal Concepts

📅 2026-09-08
📈 Citations: 1
Influential: 0
📄 PDF
🤖 AI Summary
研究通过内部概念检测癌症分割模型的错误,使用稀疏自编码器分解神经激活以识别失败案例的独特特征,从而准确预测模型错误。
📝 Abstract
Cancer segmentation models can fail silently, generating plausible but incorrect masks that risk missed findings or unnecessary biopsies. A critical question arises: Do AI models"know"when they are wrong, and if so, can we use the signal to predict their own failures? Humans do have a"Feeling of Error"(FOE): a spontaneous sense of unease that flags a potential error during thinking. We investigate whether cancer segmentation models exhibit an analogous internal signal. Unlike output-level cues (e.g., prediction confidence or uncertainty), which offer no insight into why a failure occurs and suffer from a sensitivity-quality tradeoff where high detection sensitivity could degrade overall segmentation quality. We instead propose to capture the model's FOE from its inner workings. Using mechanistic interpretability tools, specifically Sparse Autoencoders, we decompose internal neural activations into a dictionary of human-interpretable concepts and show that failure cases exhibit a distinct latent signature: fewer active concepts with lower activation magnitudes compared to successful segmentation. By training a classifier on these concept activations, we achieve accurate failure detection along with explanations for the model's mistakes. Experiments on prostate, pancreatic, and brain cancer segmentation demonstrate that our approach outperforms output-based methods in failure detection while preserving segmentation quality.
Problem

Research questions and friction points this paper is trying to address.

Cancer Segmentation
Feeling of Error
Model Failure
Innovation

Methods, ideas, or system contributions that make the work stand out.

Feeling of Error
Sparse Autoencoders
Mechanistic Interpretability
Cancer Segmentation
Failure Detection
🔎 Similar Papers
No similar papers found.
Mengmeng Ma
Mengmeng Ma
University of Delaware
Machine LearningComputer Vision
Y
Yunxiang Peng
University of Delaware, Newark, DE, USA
Tang Li
Tang Li
University of Delaware
Machine LearningComputer VisionExplainable AI
L
Lu Lin
Memorial Sloan Kettering Cancer Center, New York City, NY, USA
Binsheng Zhao
Binsheng Zhao
Memorial Sloan Kettering Cancer Center, New York City, NY, USA
O
Oguz Akin
Memorial Sloan Kettering Cancer Center, New York City, NY, USA
X
Xi Peng
University of Virginia, Charlottesville, VA, USA