From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

📅 2025-02-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Despite rapid advances, the clinical translation of generative AI in medicine remains hindered by persistent bottlenecks in transitioning from unimodal large language models to robust, clinically viable multimodal AI systems integrating medical imaging, textual narratives, and structured electronic health record data. Method: Guided by the PRISMA-ScR framework, this study conducts a systematic scoping review of 144 peer-reviewed studies published through December 2024, sourced from PubMed, IEEE Xplore, and Web of Science, with rigorous inclusion criteria. It synthesizes state-of-the-art techniques—including diffusion modeling, cross-modal alignment, explainability methods (e.g., attention visualization, saliency mapping), and clinical validation protocols. Contribution/Results: The review identifies four high-impact application domains—AI-assisted diagnosis, automated radiology/pathology reporting, generative drug discovery, and conversational clinical assistants—and delineates core challenges: data heterogeneity, limited model interpretability, ethical and regulatory uncertainties, and insufficient real-world deployment evidence. It proposes a comprehensive, trust-centered framework for responsible development and clinical integration of multimodal generative AI in healthcare.

Technology Category

Application Category

📝 Abstract
Generative artificial intelligence (AI) models, such as diffusion models and OpenAI's ChatGPT, are transforming medicine by enhancing diagnostic accuracy and automating clinical workflows. The field has advanced rapidly, evolving from text-only large language models for tasks such as clinical documentation and decision support to multimodal AI systems capable of integrating diverse data modalities, including imaging, text, and structured data, within a single model. The diverse landscape of these technologies, along with rising interest, highlights the need for a comprehensive review of their applications and potential. This scoping review explores the evolution of multimodal AI, highlighting its methods, applications, datasets, and evaluation in clinical settings. Adhering to PRISMA-ScR guidelines, we systematically queried PubMed, IEEE Xplore, and Web of Science, prioritizing recent studies published up to the end of 2024. After rigorous screening, 144 papers were included, revealing key trends and challenges in this dynamic field. Our findings underscore a shift from unimodal to multimodal approaches, driving innovations in diagnostic support, medical report generation, drug discovery, and conversational AI. However, critical challenges remain, including the integration of heterogeneous data types, improving model interpretability, addressing ethical concerns, and validating AI systems in real-world clinical settings. This review summarizes the current state of the art, identifies critical gaps, and provides insights to guide the development of scalable, trustworthy, and clinically impactful multimodal AI solutions in healthcare.
Problem

Research questions and friction points this paper is trying to address.

Exploring multimodal AI in medicine
Reviewing AI applications in clinical settings
Addressing challenges in AI integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal AI integration
Generative AI in medicine
Clinical workflow automation
L
Lukas Buess
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nßrnberg, Erlangen, Germany
Matthias Keicher
Matthias Keicher
Technische Universität Mßnchen
Nassir Navab
Nassir Navab
Professor of Computer Science, Technische Universität Mßnchen
A
Andreas Maier
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nßrnberg, Erlangen, Germany
Soroosh Tayebi Arasteh
Soroosh Tayebi Arasteh
RWTH Aachen University
Deep LearningAI in MedicineGenerative AIMedical Image Analysis