When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information

📅 2026-08-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the critical issue of overconfidence in large language models (LLMs) when confronted with ambiguous or missing clinical information, which can lead to high-risk diagnostic errors. The authors propose a dual-evaluation paradigm that integrates linguistic uncertainty prompts with answer-removal scenarios, constructing a benchmark framework based on MedMCQA and introducing multidimensional calibration metrics, including the Unsafe Confidence Error Rate (UCER). Their findings reveal that while model accuracy drops significantly under information-deficient conditions, confidence levels remain disproportionately high, resulting in a sharp increase in UCER. Furthermore, substantial variation exists across models in their ability to abstain from answering when uncertain, exposing fundamental shortcomings in current LLMs’ calibration and safety-aware decision-making capabilities.
📝 Abstract
Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their reliability under uncertainty remains poorly understood which raises critical concerns for deployment in high-stakes clinical settings. In such environments, incorrect predictions are inherently risky, but confident incorrect predictions can be particularly harmful as they may mislead clinical decision-making. In this paper, we conduct a systematic behavioral analysis of LLMs under clinical information uncertainty. We propose an evaluation framework based on the MedMCQA dataset consisting of two complementary uncertainty settings. First, we introduce linguistic uncertainty cues through prompt modifications to simulate ambiguous clinical contexts. Second, we construct an answer removal setting, wherein the correct option is deliberately excluded mandating the model to recognize insufficient information and abstain. We analyze both model accuracy and confidence behavior using multiple calibration metrics including calibration gap, Expected Calibration Error (ECE), and Unsafe Confident Error Rate (UCER) across 500 medical questions. Our results reveal a consistent failure mode, i.e., although accuracy degrades under increasing uncertainty, model confidence remains misaligned with accuracy. This leads to a substantial increase in unsafe confident errors, indicating that model confidence remains largely insensitive to clinically meaningful information loss. Furthermore, we observe significant variation across models in their ability to abstain when the correct answer is unavailable, with some models persistently producing high confidence hallucinated answers. These findings expose critical limitations in the epistemic reliability of current LLMs and highlight the need for uncertainty aware evaluation methods prior to their deployment in clinical workflows.
Problem

Research questions and friction points this paper is trying to address.

overconfidence
uncertainty
clinical reasoning
large language models
model calibration
Innovation

Methods, ideas, or system contributions that make the work stand out.

uncertainty calibration
clinical reasoning
abstention behavior
unsafe confident error
LLM reliability
M
Maryam Tahermazandarani
School of Computing, Macquarie University, Sydney, NSW 2109, Australia
Adnan Mahmood
Adnan Mahmood
School of Computing, Faculty of Science and Engineering, Macquarie University
Internet of ThingsInternet of VehiclesTrust ManagementSoftware Defined NetworkingPrivacy Preservation
F
Fahmida Islam
School of Computing, Macquarie University, Sydney, NSW 2109, Australia
Q
Quan Z. Sheng
School of Computing, Macquarie University, Sydney, NSW 2109, Australia