The widening evaluation gap in medical large language model research 2023 to 2026

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了2023至2026年间医学大语言模型评价滞后的现象,通过分析PubMed数据指出随机对照试验等设计的滞后加剧,并提出模型更新速度与临床证据生成速度不匹配的问题。
📝 Abstract
Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fourteen clinical domains, growing 45-fold; 2.5% used a randomised, controlled or prospective design. Evaluation lag, from a study's newest named model release to its own publication, widened from 1.33 to 6.08 quarters. Because discontinued models age mechanically, we benchmarked this against a counterfactual holding model composition fixed: migration to newer systems offset only 56% of the drift (95% CI 50-65). Randomised trials evaluated models a median 4.6 quarters older than other designs (P = 3 x 10^-19), yet among studies naming a model still under development no design differed from any other; 62% of randomised trials evaluated a discontinued family. Rigour and currency are in tension, and that tension reflects model selection rather than research timelines.
Problem

Research questions and friction points this paper is trying to address.

medical large language models
evaluation gap
clinical research
Innovation

Methods, ideas, or system contributions that make the work stand out.

evaluation lag
large language models
clinical research
randomised controlled trials
🔎 Similar Papers
No similar papers found.
R
Raad Bin Tareaf
XU Exponential University of Applied Sciences, Data Science and Artificial Intelligence Cluster, Marlene-Dietrich-Allee 12B, 14482 Potsdam, Germany
M
Murad Al-Rajab
College of Engineering, Abu Dhabi University, Abu Dhabi, United Arab Emirates
S
Samia Loucif
College of Technological Innovation, Zayed University, Abu Dhabi, United Arab Emirates