A primer on evaluation methods for large language models in healthcare

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了医疗领域大型语言模型的评估方法,通过研究设计原则、统计方法、能力评估和临床背景评估来确保其益处并避免潜在危害。
📝 Abstract
Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine learning for many reasons, including probabilistic and open-ended outputs, and behavior that shifts with prompt design and accumulated context. This review covers four key areas of LLM evaluation: principles of study design, statistical methods, capability evaluation and clinical context evaluation. Capability evaluation considers different benchmarks, including multiple-choice, agentic and multi-turn benchmarks, alongside operational metrics like token usage. Clinical context evaluation addresses establishing accuracy of free text outputs, such as human review and LLM-as-a-judge, and clinical trial approaches. Across sections, we describe underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question. Together, this article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs.
Problem

Research questions and friction points this paper is trying to address.

large language models
healthcare
evaluation methods
probabilistic outputs
open-ended outputs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Evaluation Methods
Healthcare Applications
Capability Evaluation
Clinical Context
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Suzannah E McKinney
Mass General Brigham AI, Mass General Brigham, Boston, MA, USA
P
Phuc Vu
Mass General Brigham AI, Mass General Brigham, Boston, MA, USA
S
Samuel A Justice
Mass General Brigham AI, Mass General Brigham, Boston, MA, USA
C
Christopher Humphries
Generative AI Laboratory, School of Informatics, University of Edinburgh, Edinburgh, UK; University of Edinburgh Centre for Cardiovascular Science, Edinburgh, UK
A
Alyssa Pradhan
Experimental Medicine Division, Nuffield Department of Medicine, University of Oxford, Oxford, UK
T
Timothy J Keyes
Department of Biomedical Data Science, Stanford University School of Medicine, Stanford University, Stanford, CA, USA
B
Bernardo C Bizzo
Mass General Brigham AI, Mass General Brigham, Boston, MA, USA; Department of Radiology, Mass General Brigham, Boston, MA, USA; Harvard Medical School, Boston, MA, USA
K
Keith J Dreyer
Mass General Brigham AI, Mass General Brigham, Boston, MA, USA; Department of Radiology, Mass General Brigham, Boston, MA, USA; Harvard Medical School, Boston, MA, USA
S
Sarah F Mercaldo
Mass General Brigham AI, Mass General Brigham, Boston, MA, USA; Department of Radiology, Mass General Brigham, Boston, MA, USA; Harvard Medical School, Boston, MA, USA
J
James M Hillis
Mass General Brigham AI, Mass General Brigham, Boston, MA, USA; Department of Neurology, Mass General Brigham, Boston, MA, USA; Harvard Medical School, Boston, MA, USA