What Carries the Signal in Pathology Foundation-Model Atlases? A Patient-Level Controlled Benchmark in Breast Cancer

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study aims to identify components within foundation models of histopathology that genuinely encode molecular pathway signals, rather than relying on cohort-level statistical associations. Evaluating eleven frozen backbone models on the TCGA-BRCA dataset at the patient level, the authors employ rigorous patient-wise GroupKFold cross-validation and permutation testing to assess predictive performance for four predefined gene pathways. Results demonstrate that embedding representations significantly outperform tissue composition features (maximum Δρ = +0.479, p ≤ 0.003), except in basal-like cases; geometric graph structures confer no additional benefit (Δρ ≈ 0); and a compact, interpretable set of 54 cellular count features nearly matches full model performance. UNI2 emerges as the top-performing model (Spearman ρ = 0.556 for immune pathways), providing the first patient-level validation of the true signal source.
📝 Abstract
Pathology foundation models are reported to encode molecular programmes in tissue morphology, but the evidence is usually a cohort-wide ranked gene list rather than a prediction for a held-out patient. We rebuild such an analysis with the patient as the unit of evidence and ask which pipeline component carries signal. Across 11 frozen backbones, four pre-specified gene programmes and 285 TCGA-BRCA patients with paired slides and RNA-seq (44 cells; GroupKFold by patient, all preprocessing fitted inside the fold), ridge regression on mean-pooled embeddings predicts held-out programme scores at Spearman rho = 0.25-0.56, UNI2 strongest on all four (immune 0.556). A matched permutation null gives raw p ~ 1e-4 at 10,000 permutations for every cell; Holm-adjusted p = 0.0044. The signal is real but not uniformly morphological. Against competing models on the same patients and folds, embeddings beat tissue composition for ER/luminal, proliferation and immune (+0.280, +0.284, +0.479; p <= 0.003) but not basal, where compartment fractions alone reach 0.469 against the embedding's 0.493 (p = 0.77). Fifty-four interpretable cell-count features come within 0.043-0.085 on every programme. The geometric machinery contributes nothing measurable, and we identify why: the geodesic graph selects neighbours by Euclidean nearest-neighbour search and only reweights edges already chosen, so the topology is Euclidean by construction (Riemannian minus Euclidean = +0.0010, 95% CI [-0.0007, +0.0029]). Applied consistently the geometry is worse (-0.0117). Ridge regression beats the graph-and-metric decoder by +0.097 (CI [+0.069, +0.127]). The driver-count metric common in this literature is near-uninformative here: 91.8% of random six-gene panels recover >=5/6 drivers.
Problem

Research questions and friction points this paper is trying to address.

pathology foundation models
molecular programmes
tissue morphology
patient-level benchmark
signal attribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

foundation model
patient-level benchmark
morphological signal
geodesic graph
ridge regression
🔎 Similar Papers
2024-08-29Medical Imaging 2025: Digital and Computational PathologyCitations: 1
C
Chimdi Walter Ndubuisi
Department of Electrical Engineering and Computer Science, University of Missouri, Columbia, MO 65211, USA