LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过模拟医患对话评估大语言模型在医疗咨询中的早期表现,发现特定指令能改善模型的建议和交接总结,但不能保证关键事实的提取。
📝 Abstract
Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24 fixed-script transcripts; two cases also used adaptive standardized-patient simulation, yielding 12 transcripts. Self-care or home-management advice before any patient answer appeared in 9 of 12 baseline case-model cells and 0 of 12 instruction cells, while structured handoff summaries appeared in 0 of 12 and 10 of 12 cells, respectively. The instruction changed sequencing and documentation, although it did not reliably ensure elicitation of decisive facts. The preformulation gap should therefore be evaluated directly through observable first-contact behavior rather than inferred from diagnostic accuracy or final-answer quality.
Problem

Research questions and friction points this paper is trying to address.

large language models
medical consultation
preformulation gap
clinical problem
Innovation

Methods, ideas, or system contributions that make the work stand out.

preformulation gap
early-stage evaluation
structured handoff summaries
🔎 Similar Papers
No similar papers found.
Y
Yining Hua
Department of Epidemiology, Harvard T.H. Chan School of Public Health, Boston, Massachusetts, USA
C
Cyrus Ayubcha
Department of Epidemiology, Harvard T.H. Chan School of Public Health, Boston, Massachusetts, USA; Harvard Medical School, Boston, Massachusetts, USA
Hongbin Na
Hongbin Na
Australian AI Institute, University of Technology Sydney / Shanghai AI Laboratory
Computational Social ScienceNarrative UnderstandingAI for Healthcare
L
Levi Lian
Raycaster, New York, New York, USA
A
Alon Gorenshtein
Department of Neurology, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, Massachusetts, USA; BRIDGE GenAI Lab, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, Massachusetts, USA
Y
Yiftach Barash
BRIDGE GenAI Lab, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, Massachusetts, USA; Department of Radiology, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, Massachusetts, USA
E
Eyal Klang
BRIDGE GenAI Lab, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, Massachusetts, USA; Department of Radiology, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, Massachusetts, USA