Specialized curricula for training vision-language models in retinal image analysis

📅 2024-07-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current vision-language models (VLMs) exhibit limited clinical utility in retinal image analysis, particularly for age-related macular degeneration (AMD) staging and referral decision-making—significantly underperforming relative to ophthalmologists. To address this gap, we propose RetinaVLM: the first AMD-specific clinical capability taxonomy, introducing a novel ophthalmology-driven VLM capability disentanglement framework and a progressive curriculum-based fine-tuning paradigm. Leveraging multi-source, image–report aligned datasets, RetinaVLM integrates curriculum-guided instruction tuning with domain-enhanced alignment techniques, adaptable to both ChatGPT-4o and medical VLM architectures. Experiments demonstrate substantial improvements over baselines: F1-scores of 0.63 for AMD staging and 0.67 for referral recommendation. In double-blind clinical evaluation, RetinaVLM achieves 64.3% report accuracy—markedly surpassing ChatGPT-4o (14.3%)—and attains performance comparable to junior ophthalmologists for the first time.

Technology Category

Application Category

📝 Abstract
Clinicians spend a significant amount of time reviewing medical images and transcribing their findings regarding patient diagnosis, referral and treatment in text form. Vision-language models (VLMs), which automatically interpret images and summarize their findings as text, have enormous potential to alleviate clinical workloads and increase patient access to high-quality medical care. While foundational models have stirred considerable interest in the medical community, it is unclear whether their general capabilities translate to real-world clinical utility. In this work, we demonstrate that OpenAI's ChatGPT-4o model, in addition to two foundation VLMs designed for medical use, markedly underperform compared to practicing ophthalmologists on specialist tasks crucial to the care of patients with age-related macular degeneration (AMD). To address this, we initially identified the essential capabilities required for image-based clinical decision-making, and then developed a curriculum to selectively train VLMs in these skills. The resulting model, RetinaVLM, can be instructed to write reports that significantly outperform those written by leading foundation medical VLMs and ChatGPT-4o in disease staging (F1 score of 0.63 vs. 0.33) and patient referral (0.67 vs. 0.50), and approaches the diagnostic performance of junior ophthalmologists (who achieve 0.77 and 0.78 on the respective tasks). Furthermore, in a single-blind reader study two senior ophthalmologists with up to 32 years of experience found RetinaVLM's reports were found to be substantially more accurate than those by ChatGPT-4o (64.3% vs. 14.3%). These results reinforce that our curriculum-based approach provides a blueprint towards specializing foundation medical VLMs for real-world clinical tasks.
Problem

Research questions and friction points this paper is trying to address.

Improving vision-language models for retinal image analysis
Enhancing clinical decision-making in age-related macular degeneration
Developing specialized curricula for medical vision-language training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Curriculum-based VLM specialization
Enhanced retinal image analysis
Outperforms general medical VLMs
💼 Related Jobs
No related jobs found.
Imperial College London | University of Southampton | Moorfields Eye Hospital NHS Foundation Trust | Medical University of Vienna | Technical University of Munich | Institute of Molecular and Clinical Ophthalmology Basel | University of Basel
R
Robbie Holland
Biomedical Image Analysis, Department of Computing, Imperial College London, United Kingdom
T
Thomas R. P. Taylor
Clinical and Experimental Sciences, Faculty of Medicine, University of Southampton, United Kingdom
C
Christopher Holmes
Moorfields Eye Hospital NHS Foundation Trust, London, United Kingdom
S
Sophie Riedl
Laboratory for Ophthalmic Image Analysis, Medical University of Vienna, Austria
J
Julia Mai
Laboratory for Ophthalmic Image Analysis, Medical University of Vienna, Austria
M
Maria Patsiamanidi
Clinical and Experimental Sciences, Faculty of Medicine, University of Southampton, United Kingdom
D
Dimitra Mitsopoulou
Clinical and Experimental Sciences, Faculty of Medicine, University of Southampton, United Kingdom
P
Paul Hager
Institute for AI in Healthcare and Medicine, Klinikum rechts der Isar, Technical University of Munich, Germany
P
Philip Muller
Institute for AI in Healthcare and Medicine, Klinikum rechts der Isar, Technical University of Munich, Germany
H
H. Scholl
Institute of Molecular and Clinical Ophthalmology Basel, Switzerland; Department of Ophthalmology, University of Basel, Switzerland; Department of Clinical Pharmacology, Medical University of Vienna, Austria
H
Hrvoje Bogunovi'c
Laboratory for Ophthalmic Image Analysis, Medical University of Vienna, Austria
U
U. Schmidt-Erfurth
Laboratory for Ophthalmic Image Analysis, Medical University of Vienna, Austria
Daniel Rueckert
Daniel Rueckert
Technical University of Munich and Imperial College London
Machine LearningMedical Image ComputingBiomedical Image AnalysisComputer Vision
S
S. Sivaprasad
Moorfields Eye Hospital NHS Foundation Trust, London, United Kingdom
A
A. Lotery
Clinical and Experimental Sciences, Faculty of Medicine, University of Southampton, United Kingdom
M
M. Menten
Biomedical Image Analysis, Department of Computing, Imperial College London, United Kingdom; Institute for AI in Healthcare and Medicine, Klinikum rechts der Isar, Technical University of Munich, Germany