Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions

๐Ÿ“… 2026-08-01
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the absence of confidence metrics in existing medical vision-language models for chest X-ray diagnosis and the lack of systematic evaluation regarding their alignment with local institutional reference standards. We present the first quantitative assessment of cross-institutional consistency between generative vision-language model outputs and local standards, introducing multiple few-shot estimation methods leveraging limited local labels. Rigorous leave-one-institution-out validation was conducted across three institutions encompassing over 345,000 predictions. Results demonstrate that Beta-Binomial empirical Bayes and target-domain logistic regression achieve the best performance (Brier score โ‰ˆ 0.085), significantly outperforming a cross-institutional aggregation baselineโ€”though this advantage is institution-specific. Notably, nominal 95% prediction intervals exhibit only 87.0% empirical coverage, underscoring the necessity of institution- and interface-specific calibration for reliable consistency evaluation.
๐Ÿ“ Abstract
Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far to trust an individual judgment. Whether agreement with an institution's reference standard transfers across sites, findings, prediction directions and question formats is largely unmeasured. We evaluated three generative vision-language models on three institutional chest-radiograph corpora and six findings under two elicitation protocols, comprising more than 345,000 finding-level predictions, and estimated finding-by-direction reference agreement at a receiving institution from a small budget of local labels. Estimation strategies were then stress-tested under repeated strict institution-held-out evaluation. Under evaluation excluding the receiving institution from development entirely, adaptive selection among the seven estimators that design admits did not improve on simple fixed alternatives: it achieved a mean Brier score of 0.1083, against 0.0853 for always using a Beta-Binomial empirical-Bayes estimator and 0.0855 for a target-only logistic model. Those two differ by 0.0003, less than this family's own sensitivity to a change of solver version, and each leads in about half the settings, so no default can be recommended. Their advantage over estimators pooling across institutions was concentrated at one site and not confirmatory once clustered by institution, and a plug-in empirical-Bayes posterior-predictive count interval at a nominal 95% level covered 87.0%, less at the hardest institution. Reference agreement therefore has to be re-evaluated per site and per interface; these results concern agreement with institutional labels, not clinical correctness.
Problem

Research questions and friction points this paper is trying to address.

medical vision-language models
chest radiographs
reference agreement
institutional transferability
confidence estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language models
reference agreement
cross-institutional evaluation
empirical Bayes estimation
medical imaging
๐Ÿ’ผ Related Jobs
No related jobs found.
P
Pengyang Yu
School of Computer Science, University College Dublin, Dublin, Ireland
Y
Yiou Wang
Department of Medical Imaging, The Third Affiliated Hospital of Southern Medical University, Guangzhou, China
Z
Zhongping Dong
School of Computer Science, University College Dublin, Dublin, Ireland
S
Sahraoui Dhelim
Dublin City University, Dublin, Ireland
Chun-Mei Feng
Chun-Mei Feng
Assistant Professor/Ad Astra Fellow, University College Dublin, Ireland
AI for HealthCareMulti-modal LearningFederated Learning
M
M. Tahar Kechadi
School of Computer Science, University College Dublin, Dublin, Ireland