Explanation Stability of Test-Time Adaptation in Computational Pathology: A Large-Scale Benchmark

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the underexplored trade-off between accuracy and explanation stability in test-time adaptation (TTA) for computational pathology. While TTA can enhance model accuracy, it may compromise the stability of model explanations, thereby undermining clinical trustworthiness. The authors present the first large-scale systematic evaluation of 17 TTA methods across the Camelyon17 and NCT-CRC-HE datasets, examining their impact on explanation stability using four attribution techniques and encompassing convolutional networks, Vision Transformers, and foundation models in pathology—totaling 2,958 experiments. Their findings reveal that explanation stability is decoupled from predictive accuracy and should be treated as an independent reliability metric for TTA. Notably, TTA strategies that freeze the backbone yield the most stable explanations, whereas continual adaptation approaches like CoTTA induce significant explanation drift. Convolutional architectures exhibit greater sensitivity to TTA-induced instability than Transformers. The complete benchmark and evaluation protocol are publicly released.
📝 Abstract
Test-time adaptation (TTA) has become a practical way to adapt deployed models to unlabeled target data, a setting that is especially relevant in computational pathology where staining, scanner, and cohort shifts are routine. While most TTA methods are evaluated by their effect on accuracy, clinical use also depends on whether the model's explanations remain reliable after adaptation. In this paper, we take a closer look at this largely unmeasured effect. We study explanation stability under TTA across two histopathology benchmarks, Camelyon17 and NCT CRC-HE, using five architectures ranging from convolutional networks to vision transformers and a pathology foundation model, seventeen TTA methods, and four attribution families. Across 2,958 adaptation runs, we observe a clear and systematic pattern: TTA methods differ sharply in how much they move model explanations, with frozen-backbone methods leaving attributions almost unchanged and continual methods such as CoTTA and RoTTA causing the largest drift. This effect is not uniform. Convolutional networks are substantially more sensitive than transformer and foundation-model backbones, and explanation drift increases with adaptation strength while remaining largely insensitive to batch size. Surprisingly, explanation stability is only weakly coupled to adaptation quality. Some methods preserve explanations almost perfectly while degrading calibration or accuracy, producing silent failures that would be missed by accuracy-only or explanation-only evaluation. These findings show that explanation stability is a distinct reliability axis for TTA in computational pathology. We release the metric, protocol, and full benchmark to support future work on adaptation methods that are not only accurate, but also stable and clinically auditable. Code: https://github.com/bahumanyarg11/tta-explanation-stability-pipeline
Problem

Research questions and friction points this paper is trying to address.

Test-time adaptation
Explanation stability
Computational pathology
Model interpretability
Domain shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Adaptation
Explanation Stability
Computational Pathology
Attribution Drift
Model Reliability
🔎 Similar Papers
No similar papers found.
R
R. G. Bahumanya
Department of Computer Science and Engineering, R.V. College of Engineering, Bengaluru, India
H
Harshith V. M.
Department of Computer Science and Engineering, R.V. College of Engineering, Bengaluru, India
S
Shreyank N. Gowda
School of Computer Science, University of Nottingham, Nottingham, United Kingdom
A
Anala M. R.
Department of Computer Science and Engineering, R.V. College of Engineering, Bengaluru, India