đ¤ AI Summary
Despite rapid advances, the clinical translation of generative AI in medicine remains hindered by persistent bottlenecks in transitioning from unimodal large language models to robust, clinically viable multimodal AI systems integrating medical imaging, textual narratives, and structured electronic health record data. Method: Guided by the PRISMA-ScR framework, this study conducts a systematic scoping review of 144 peer-reviewed studies published through December 2024, sourced from PubMed, IEEE Xplore, and Web of Science, with rigorous inclusion criteria. It synthesizes state-of-the-art techniquesâincluding diffusion modeling, cross-modal alignment, explainability methods (e.g., attention visualization, saliency mapping), and clinical validation protocols. Contribution/Results: The review identifies four high-impact application domainsâAI-assisted diagnosis, automated radiology/pathology reporting, generative drug discovery, and conversational clinical assistantsâand delineates core challenges: data heterogeneity, limited model interpretability, ethical and regulatory uncertainties, and insufficient real-world deployment evidence. It proposes a comprehensive, trust-centered framework for responsible development and clinical integration of multimodal generative AI in healthcare.
đ Abstract
Generative artificial intelligence (AI) models, such as diffusion models and OpenAI's ChatGPT, are transforming medicine by enhancing diagnostic accuracy and automating clinical workflows. The field has advanced rapidly, evolving from text-only large language models for tasks such as clinical documentation and decision support to multimodal AI systems capable of integrating diverse data modalities, including imaging, text, and structured data, within a single model. The diverse landscape of these technologies, along with rising interest, highlights the need for a comprehensive review of their applications and potential. This scoping review explores the evolution of multimodal AI, highlighting its methods, applications, datasets, and evaluation in clinical settings. Adhering to PRISMA-ScR guidelines, we systematically queried PubMed, IEEE Xplore, and Web of Science, prioritizing recent studies published up to the end of 2024. After rigorous screening, 144 papers were included, revealing key trends and challenges in this dynamic field. Our findings underscore a shift from unimodal to multimodal approaches, driving innovations in diagnostic support, medical report generation, drug discovery, and conversational AI. However, critical challenges remain, including the integration of heterogeneous data types, improving model interpretability, addressing ethical concerns, and validating AI systems in real-world clinical settings. This review summarizes the current state of the art, identifies critical gaps, and provides insights to guide the development of scalable, trustworthy, and clinically impactful multimodal AI solutions in healthcare.