A visual large language foundational model for medical image recognition using clinician-oriented social media

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决医学图像识别中临床推理和图文对齐数据稀缺问题,通过结合高级大语言模型与临床医生验证,构建了包含超过一百万VQA对的ThoughtMed-1M数据集。
📝 Abstract
Large language models (LLMs) have demonstrated strong capabilities across diverse domains, showing considerable potential in medicine. However, their application in medical settings remains limited by the scarcity of visual question answering (VQA) datasets that capture clinical reasoning and explicit image-text alignment. Here, we leverage de-identified medical images and expert commentaries shared on clinician-oriented social media. By combining an advanced LLM with clinician-in-the-loop verification, we established a rigorous pipeline to construct ThoughtMed-1M, a long-form medical VQA dataset containing over one million VQA pairs and designed to capture structured clinical logic and medical image-text alignment. To demonstrate its utility, we developed a FOundational LLM Trained on ThoughtMed-1M (FOLTMed). FOLTMed achieved state-of-the-art performance across 42 medical VQA benchmark datasets, with a macro accuracy of 85.4%, and generated more clinically coherent responses on the ThoughtMed-1M test set. It outperformed state-of-the-art models by 3--5% across factuality and similarity metrics, highlighting a scalable paradigm for advancing research on clinically grounded multimodal LLMs.
Problem

Research questions and friction points this paper is trying to address.

large language models
medical image recognition
visual question answering
clinical reasoning
image-text alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual large language model
medical VQA dataset
clinician-in-the-loop verification
clinical reasoning
💼 Related Jobs
No related jobs found.
L
Lingxuan Hou
College of Biomedical Engineering, Sichuan University, Chengdu, Sichuan, China
Y
Yuhua Xie
College of Biomedical Engineering, Sichuan University, Chengdu, Sichuan, China
Y
Yue Hu
Guang’anmen Hospital, China Academy of Chinese Medical Sciences, Beijing, China
Yan Zhuang
Yan Zhuang
University of Science and Technology of China
computerized adaptive testingAI education
J
Junqi Li
Jincheng General Hospital, Jincheng, Shanxi, China
C
Chengzhi Xia
College of Biomedical Engineering, Sichuan University, Chengdu, Sichuan, China
B
Binh Phu Nguyen
School of Mathematics and Statistics, Victoria University of Wellington, Wellington, New Zealand
Abubakar Siddique
Abubakar Siddique
School of Engineering, Computer & Mathematical Sciences, Auckland University of Technology, Auckland, New Zealand
Minh Nguyen
Minh Nguyen
Head of Department of Computer and Information Sciences, Professor, AUT
Computer VisionVirtual Reality and Related SimulationComputer-Human InteractionKnowledge
Y
Yao Hou
College of Biomedical Engineering, Sichuan University, Chengdu, Sichuan, China
Y
Yanju Bao
Guang’anmen Hospital, China Academy of Chinese Medical Sciences, Beijing, China
K
Kexin Liu
Guang’anmen Hospital, China Academy of Chinese Medical Sciences, Beijing, China
K
Ke Chen
College of Biomedical Engineering, Sichuan University, Chengdu, Sichuan, China
J
Jianjun Sun
College of Biomedical Engineering, Sichuan University, Chengdu, Sichuan, China
Z
Zeqi Li
Jincheng General Hospital, Jincheng, Shanxi, China
T
Trung Nguyen
School of Engineering, Computer & Mathematical Sciences, Auckland University of Technology, Auckland, New Zealand
J
Jiangli Lin
College of Biomedical Engineering, Sichuan University, Chengdu, Sichuan, China