Rethinking How We Evaluate Methodological Progress in Health AI

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过重新实现12种算法并使用共享框架在两个临床数据集上评估,探讨了健康AI中方法学进展的可重复性和临床任务定义问题。
📝 Abstract
Methodological progress in artificial intelligence (AI) for electronic health records (EHRs) depends on our ability to determine which algorithms work better, and under which conditions. However, such progress is thought to be hindered by difficulties in reproducibility and in defining clinically meaningful evaluation tasks. We empirically study these barriers by re-implementing 12 historical and recent algorithms within a shared evaluation framework and evaluating them on two clinical datasets, MIMIC-IV and NWICU. We compare two complementary task families: expert-authored clinically meaningful tasks and generated tasks defined from randomly sampled event codes and prediction horizons. We ask whether relative algorithms comparisons transfer across task families and datasets, whether residual task heterogeneity contains useful methodological structure, and what a controlled comparison reveals about progress over the last decade. We find that aggregate pairwise comparisons transfer strongly across evaluation settings, including from randomly generated tasks to clinically meaningful tasks and across datasets. At the same time, clinically meaningful tasks exhibit greater task-method interaction, providing preliminary evidence that task properties can help explain when particular modeling choices are advantageous. Finally, newer algorithms do not consistently outperform earlier approaches: gradient-boosted trees remain highly competitive when paired with a modern, wide and sparse representation of the EHR. Together, these results suggest that useful methodological knowledge may require less task engineering than commonly assumed, while highlighting the importance of understanding the structured heterogeneity that remains across tasks and methods.
Problem

Research questions and friction points this paper is trying to address.

methodological progress
electronic health records
reproducibility
clinically meaningful evaluation tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

algorithm re-implementation
shared evaluation framework
task-method interaction
gradient-boosted trees
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
F
Florent Pollet
Department of Biomedical Informatics, Columbia University
M
Matthew McDermott
Department of Biomedical Informatics, Columbia University