MedReaMM: Evaluating Large Multimodal Models on Expert-Level Clinical Diagnostic Synthesis

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决临床诊断综合能力评估问题,研究通过MedReaMM基准测试大型多模态模型整合病史与医学影像进行准确诊断的能力。
📝 Abstract
The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest. However, existing benchmarks largely focus on textual reasoning or isolated visual question-answering (VQA) tasks, lacking holistic integration of clinical narratives and medical imaging, and thus failing to assess the multimodal diagnostic synthesis capability central to expert clinical judgment. To bridge this gap, we introduce MedReaMM, a benchmark specifically designed to evaluate models' ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm. Constructed from case reports sourced from top-tier medical journals and curated clinical case databases, MedReaMM comprises 625 expert-validated cases with an average of 2.79 medical images per case and a total of 1,042 standardized diagnoses annotated with ICD-11 codes. These cases predominantly represent rare, atypical, or multi-system presentations that demand expert-level evidence integration beyond routine pattern recognition. We evaluate 23 Large Multimodal Models (LMMs) and find that most achieve diagnostic accuracy scores below 50%, underscoring a substantial gap in multimodal diagnostic synthesis capability. Further analysis reveals that medical knowledge proficiency, medical image understanding, and evidence integration are all highly correlated with diagnostic performance.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
multimodal diagnostic synthesis
clinical narratives
medical imaging
expert clinical judgment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Multimodal Models
Clinical Diagnostic Synthesis
Multimodal Integration
Medical Imaging
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Lai Wei
Lai Wei
Huazhong University of Science and technology
Multimodal Language ModelsHealthcare
Yuchao Chen
Yuchao Chen
WellSIM Biomedical Technologies
MicrofluidicsBiomedical deviceExtracellular vesicles
Z
Zhenbiao Cao
Huazhong University of Science and Technology
X
Xiaojin Zhang
Huazhong University of Science and Technology
Z
Zhongyu Wei
Fudan University
B
Bangting Wang
The First Affiliated Hospital with Nanjing Medical University and Jiangsu Province Hospital
W
Wei Chen
Huazhong University of Science and Technology
Xiang Bai
Xiang Bai
Huazhong University of Science and Technology (HUST)
Computer VisionOCR