Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

๐Ÿ“… 2026-08-11
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the tendency of large language models to rely on superficial heuristics and generate overconfident responses in medical hypothesis verification, which compromises safety and reliability. To mitigate this issue, the authors propose a two-stage verification framework that treats a modelโ€™s active abstention as a signal of uncertainty, triggering ontology-grounded selective reasoning refinement. Notably, this approach achieves knowledge graphโ€“level performance without explicitly constructing a knowledge graph. Evaluated on MedReason and MedQA benchmarks, the method substantially improves performance, attaining a question-level accuracy of 92.5% (+9.6%) and a hypothesis-level accuracy of 96.2% (+4.2%), effectively balancing coverage and precision.
๐Ÿ“ Abstract
Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models to abstain when uncertain improves reliability but introduces a coverage accuracy tradeoff. We propose a two-stage framework for medical hypothesis verification in multiple-choice settings that manages this tradeoff through targeted ontology grounding, applied only when the model abstains. We show that abstention is not random but reflects genuine uncertainty, with abstained predictions associated with lower confidence. Across two frontier models (GPT-5.5, accessed via the Azure OpenAI API, and DeepSeek-R1), the proposed framework improves question-level accuracy by 9.6 percentage points (82.9% to 92.5%) and hypothesis-level accuracy by 4.2 percentage points (92.0% to 96.2%). Our experiments conducted on MedReason and MedQA show that abstention can be repurposed as a control signal for selective reasoning refinement, achieving knowledge-graph-level performance without explicit knowledge graph construction.
Problem

Research questions and friction points this paper is trying to address.

LLM verification
medical reasoning
abstention
uncertainty
coverage-accuracy tradeoff
Innovation

Methods, ideas, or system contributions that make the work stand out.

abstention
ontology grounding
selective refinement
uncertainty-aware verification
medical reasoning
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
U
Uma Ranjan
Indian Institute of Technology Jammu
K
Kunal Tilaganji
Microsoft Research
A
Aditya Koul
Indian Institute of Technology Jammu
A
Anurag Mahipal
Indian Institute of Technology Jammu
D
Dashpreet Singh
Indian Institute of Technology Jammu
H
Hriday Rana
Indian Institute of Technology Jammu
M
Manan Jain
Indian Institute of Technology Jammu
Sidharth Gupta
Sidharth Gupta
Indian Institute of Technology Jammu
A
Ajo Babu George
SCB Dental College and Hospital, Cuttack
V
Vineeth Balasubramanian
Microsoft Research
Nagarajan Natarajan
Nagarajan Natarajan
Researcher, Microsoft Research India
Machine LearningAI for Code
Amit Sharma
Amit Sharma
Principal Researcher, Microsoft Research
Causal machine learningReasoningTrustworthy ML