ObGynLongBench: Revealing the Evidence-to-EHR Gap in Longitudinal EHR Decision-Making

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过构建ObGynLongBench基准,评估了大语言模型在真实纵向电子健康记录上的临床决策能力,揭示了证据到EHR的差距,并指出有效利用患者特定证据是关键挑战。
📝 Abstract
The application of large language models (LLMs) to personalized medical assistants has garnered growing interest. However, existing medical benchmarks largely rely on static question answering with pre-selected evidence, leaving unclear whether LLMs can make reliable clinical decisions from real longitudinal electronic health records (EHRs). To bridge this gap, we introduce ObGynLongBench, a rule-grounded long-context EHR benchmark for obstetric and gynecologic decision-making, comprising 1,500 clinical decision-point cases from 976 real pregnancy EHR histories and traceable rules. Each case is anchored to a patient, a pregnancy-timeline point, and a pre-decision information boundary, enabling Evidence-only, Visit-level EHR, and History-level EHR evaluation. Evaluating 17 LLMs reveals a substantial Evidence-to-EHR Gap: models perform well when evidence is directly provided, but accuracy drops when evidence must be extracted from same-day records or full pre-decision EHR histories. Further analyses identify evidence utilization as a key bottleneck: performance decreases with longer EHR contexts and more complex evidence requirements, and earlier failures often predict later failures within the same patient history. Finally, active-search agents perform best among EHR access strategies, highlighting patient-specific evidence utilization as a central challenge for reliable personalized medical assistants. Resources are available at https://github.com/xiangjun2003/ObgynLongbench.
Problem

Research questions and friction points this paper is trying to address.

large language models
electronic health records
clinical decision-making
benchmark
evidence utilization
Innovation

Methods, ideas, or system contributions that make the work stand out.

ObGynLongBench
Evidence-to-EHR Gap
large language models (LLMs)
longitudinal EHR decision-making
active-search agents
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jun Xiang
School of Data Science, Fudan University, China
Z
Zhijie Bao
School of Data Science, Fudan University, China; Shanghai Innovation Institute, China
Rong Hu
Rong Hu
Hunan University
K
Kaizhou Qin
Obstetrics & Gynecology Hospital of Fudan University, China
W
Wei Chen
School of Software Engineering, Huazhong University of Science and Technology, China
Z
Zhongyu Wei
School of Data Science, Fudan University, China; Shanghai Innovation Institute, China