Rethinking Correctness for Uncertainty Estimation in Clinical Prediction with Vision-Language Models

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过提出一个评估框架,解决临床预测中不确定性估计的准确性问题,以提高模型可靠性。
📝 Abstract
Vision-language models are increasingly explored for clinical prediction from electronic health records and medical images, where identifying unreliable predictions is important for safe deployment. Uncertainty estimation (UE) enables detecting such predictions, but its evaluation depends on a correctness criterion that determines whether each model output is correct. If this criterion disagrees with human judgement or distorts downstream UE performance, conclusions about model reliability can be misleading. We introduce a two-axis framework that evaluates correctness criteria by their agreement with human judgements and fidelity to human-referenced UE performance. We assess eight criteria across three clinical prediction tasks and three models using 450 predictions annotated by two reviewers. Across the audited tasks, canonical exact matching (EM) achieved the highest observed human agreement and lowest UE distortion, while the BERT-based matching (BEM) and LLM-judge also showed strong human agreement. Across four UE methods and 23,254 clinical predictions, criterion choice changed error-detection AUROC by up to 0.146 and reversed the relative ranking of UE methods. The LLM-judge also selectively accepted invalid or uncertain outputs, accepting 16 of 30 such human-identified errors. These results demonstrate that correctness assessment is an integral component of clinical UE evaluation and should be validated before UE methods are compared.
Problem

Research questions and friction points this paper is trying to address.

Uncertainty Estimation
Clinical Prediction
Correctness Criterion
Human Judgement
Model Reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uncertainty Estimation
Correctness Criteria
Human Agreement
Clinical Prediction
Vision-Language Models
🔎 Similar Papers
No similar papers found.