Scaling Clinical Judgment to Evaluate Medical AI

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过开发PrecepTron模型和GRAND-ROUNDS基准,解决了大规模评估医学AI临床推理能力的问题,采用低秩适应方法对少量医生示例进行微调。
📝 Abstract
Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific questions investigated and makes it unclear whether findings would be reproduced with a different set of evaluators. To more rigorously and scalably study clinical reasoning in AI models, here we introduce PrecepTron, an LLM fine-tuned for physician-level evaluation of open-ended responses. PrecepTron was trained using low-rank adaptation (LoRA) of a 32-billion-parameter model on a small number of physician examples. We also release GRAND-ROUNDS, a new large-scale physician-annotated benchmark of 9,217 scores by 11 physicians across seven studies. We show that frontier LLMs in typical "LLM-as-a-judge" approaches often disagree with physicians and with each other, but fine-tuning PrecepTron on a small number of cases enables physician-level consistent scoring across tasks. We use PrecepTron to reproduce headline findings from five influential studies assessing LLMs for clinical care in JAMA, Science, and Nature Medicine without new human grading. Using PrecepTron, we then pose new questions about how LLMs reason in medicine that would have been infeasible with human grading alone, including measuring the diagnostic accuracy of frontier LLMs when clinical cases are provided piecemeal, even token by token. Together, PrecepTron and GRAND-ROUNDS provide a foundation for reproducible, large-scale study of how LLMs reason in medicine. All code, data, and labels are made freely available for researchers.
Problem

Research questions and friction points this paper is trying to address.

clinical reasoning
large language models
scalability
reproducibility
physician evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

PrecepTron
low-rank adaptation (LoRA)
GRAND-ROUNDS
physician-level evaluation
🔎 Similar Papers
No similar papers found.
T
Thomas A. Buckley
Department of Biomedical Informatics, Harvard Medical School, Boston, MA
Z
Zahir Kanjee
Department of Medicine, Beth Israel Deaconess Medical Center, Boston, MA
P
Peter G. Brodeur
Department of Medicine, Beth Israel Deaconess Medical Center, Boston, MA
B
Byron Crowe
Stanford Division of Hospital Medicine, Stanford University, Stanford, CA
A
Anthony M. Pettinato
Department of Medicine, Beth Israel Deaconess Medical Center, Boston, MA
A
Aashna P. Shah
Department of Biomedical Informatics, Harvard Medical School, Boston, MA
A
Adrian D. Haimovich
Department of Emergency Medicine, Beth Israel Deaconess Medical Center, Boston, MA
L
Liam G. McCoy
Department of Medicine, Beth Israel Deaconess Medical Center, Boston, MA; Division of Neurology, University of Alberta, Edmonton, Canada; Institute for Medical Engineering and Science, MIT, Cambridge, MA
D
Daniel Restrepo
Department of Medicine, Massachusetts General Hospital, Boston, MA
E
Ethan Goh
Stanford Division of Computational Medicine, Stanford University, Stanford, CA; Stanford Clinical Excellence Research Center, Stanford University, Stanford, CA
J
Jonathan H. Chen
Stanford Division of Hospital Medicine, Stanford University, Stanford, CA; Stanford Division of Computational Medicine, Stanford University, Stanford, CA; Stanford Clinical Excellence Research Center, Stanford University, Stanford, CA
L
Laura Zwaan
Institute of Medical Education Research, Erasmus Medical Center, Rotterdam, The Netherlands
K
Katherine E. Goodman
Department of Epidemiology and Public Health, University of Maryland School of Medicine, Baltimore, MD; University of Maryland Institute for Health Computing, Bethesda, MD
Daniel J. Morgan
Daniel J. Morgan
Department of Epidemiology and Public Health, University of Maryland School of Medicine, Baltimore, MD; VA Maryland Healthcare System, Baltimore, MD
R
Raja-Elie E. Abdulnour
Department of Pulmonary and Critical Care Medicine, Brigham and Women’s Hospital, Boston, MA
Adam Rodman
Adam Rodman
Assistant Professor of Medicine, Harvard Medical School
Clinical reasoningAIdigital educationmedical history
A
Arjun K. Manrai
Department of Biomedical Informatics, Harvard Medical School, Boston, MA