CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the gap between strong performance of large language models on medical benchmarks and their limited capacity for defensible reasoning in real-world clinical settings using heterogeneous, longitudinal electronic health records (EHRs). To bridge this gap, the authors introduce CliniCARE-Bench—the first benchmark designed for clinical auditing—comprising 25 physician-validated scenarios and 750 real-world cases derived from MIMIC-IV. Agents must integrate structured and unstructured EHR data within a constrained tool environment to issue one of four policy-compliant rulings while fully tracing their investigative process. The framework uniquely unifies longitudinal EHR interrogation, evidence provenance, policy adherence, procedural compliance, and a calibrated abstention mechanism that distinguishes “missing data” from “medical ambiguity.” Reference rulings are established via multi-model arbitration and clinical committee calibration. Evaluation across 16 systems shows four-class accuracy of 65.3%–76.1%, yet defect-free accuracy—excluding shortcut-based responses—drops markedly by 4.8–14.8 percentage points, revealing substantial overestimation by conventional metrics.
📝 Abstract
Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably. We introduce CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR), a benchmark for retrospective clinical audit: 25 clinician-validated scenarios instantiated as 750 patient-specific cases over real-patient-derived MIMIC-IV data. Systems investigate each case through a governed, logged tool environment for record retrieval, computation, and policy access, and return one of four verdicts---Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous---the last two separating missing evidence from residual medical ambiguity. Beyond verdict accuracy, we score patient-evidence and policy grounding, process adherence, calibrated abstention, reliability, and efficiency against case-level reference verdicts produced by independent multi-model adjudication and calibrated against Clinical Board review. Every retrieval, computation, and report is replayable, so the investigation trace is inspectable and scorable. To our knowledge, CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework. Across 16 agentic systems, four-way accuracy spans 65.3-76.1%, but raw accuracy overstates investigation quality. Defect-free accuracy, which credits a verdict only when correct and free of prohibited shortcuts, is 4.8-14.8 points lower and reorders the leaderboard.
Problem

Research questions and friction points this paper is trying to address.

clinical reasoning
electronic health records
medical audit
evidence grounding
calibrated abstention
Innovation

Methods, ideas, or system contributions that make the work stand out.

CliniCARE-Bench
clinical reasoning
electronic health records (EHR)
calibrated abstention
evidence grounding
🔎 Similar Papers
No similar papers found.
Veronica Chatrath
Veronica Chatrath
Technical Program Manager | Vector Institute
B
Bryan Zhu
Scale AI
G
George Pu
Scale AI
Jingxuan Fan
Jingxuan Fan
Harvard
systems neuroscienceLLM
Apaar Shanker
Apaar Shanker
PhD student at Georgia Institute of Technology
Machine LearningMaterials InformaticsMaterials GenomicsComputational Materials Science
V
Varun Ursekar
Scale AI
A
Anahita Sharma
Scale AI
J
Jason Qin
Scale AI
Keqi Han
Keqi Han
Emory University
Graph MiningNetwork Analysis
S
Soham Dinesh Tiwari
Scale AI
Soham Dan
Soham Dan
Senior Research Scientist, Microsoft
Large Language ModelsNatural Language ProcessingMachine LearningArtificial Intelligence
V
Vijay Kalmath
Scale AI
Y
Yuan (Christy) Li
Scale AI
Daniel Yue Zhang
Daniel Yue Zhang
Amazon AGI, University of Notre Dame
Natural Language UnderstandingMisinformation Detection Edge Computing Human-Cyber-Physical Systems
Chenguang Wang
Chenguang Wang
UC Santa Cruz
Z
Zainab Doctor
Scale AI
Z
Zhijun Yin
Scale AI, Vanderbilt University Medical Center
N
Nigam H. Shah
Department of Medicine, Stanford School of Medicine, Technology and Digital Solutions, Stanford Health Care, Clinical Excellence Research Center, Stanford School of Medicine
Y
Yuan Xue
Scale AI