Volumetric Radiology AI in the Era of Multimodal Large Language Models

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了利用多模态大语言模型解决体积放射学中三维信息保留与理解的问题,通过体积基础模型、语言对齐等方法提升AI在临床应用中的可靠性。
📝 Abstract
Advances in multimodal large language models (MLLMs) are extending radiological artificial intelligence (AI) beyond task-specific image analysis toward multimodal understanding and reasoning. Volumetric radiology, however, presents a fundamental representational mismatch: clinical interpretation often requires full-volume spatial context and acquisition-dependent quantitative information, whereas current MLLMs are commonly conditioned on selected two-dimensional (2D) images, compressed visual representations, or report-derived text. Reliable volumetric radiology AI therefore requires representations that preserve task-relevant three-dimensional (3D) information and systems that can access, verify, and integrate this information across clinical workflows. In this Review, we examine more than 200 publications through July 2026. We organize the literature around volumetric representation and multimodal understanding at the model level, agentic orchestration at the system level, and their links to clinical applications and evaluation. We review volumetric foundation models, language alignment and compression strategies, and agentic systems that extend MLLMs through planning, tools, memory, and workflow interaction. We distinguish settings in which selected 2D views or report-mediated reasoning may suffice from those that warrant native volumetric modeling. We also introduce a Claim-Design-Validation framework to assess whether technical, workflow, and clinical claims are matched by appropriate design and validation. Across the literature, native volumetric modeling and agentic capabilities depend on the spatial, quantitative, contextual, and workflow requirements of the intended task. Clinical credibility requires faithful volumetric representation, traceable system behavior, claim-aligned validation, and clearly defined human oversight in realistic workflows.
Problem

Research questions and friction points this paper is trying to address.

volumetric radiology
multimodal large language models
spatial context
quantitative information
clinical workflows
Innovation

Methods, ideas, or system contributions that make the work stand out.

volumetric radiology
multimodal large language models
3D information preservation
agentic orchestration
Claim-Design-Validation framework
Zanting Ye
Zanting Ye
Southern Medical University
Deep learningMedical Imgae analysisVLM
Shengyuan Liu
Shengyuan Liu
The Chinese University of Hong Kong; CASIA
Multimodal LearningGenerative modelsAI for HealthcareRadiomics
X
Xin Liu
Southern Medical University
Chenhui Wang
Chenhui Wang
PhD Candidate, Fudan University
AI for NeuroscienceComputer Vision
Z
Zhisong Wang
Northwestern Polytechnical University
J
Jiashuai Liu
Xi’an Jiaotong University
Z
Zipei Wang
Institute of Automation, Chinese Academy of Sciences
Cheng Wang
Cheng Wang
City University of Hong Kong
Nanophotonics
W
Wentao Pan
The Chinese University of Hong Kong
Mengjie Fang
Mengjie Fang
Beihang University
medical image analysisradiomicspattern recognitionmachine learning
Di Dong
Di Dong
Institute of Automation, Chinese Academy of Sciences
Radiomicsdeep learninggastric cancerhead and neck cancernasopharyngeal cancer
M
Mohammad Salmanpour
University of British Columbia
Arman Rahmim
Arman Rahmim
Professor of Radiology, Physics and Biomedical Engineering, University of British Columbia
computational imagingmolecular imagingpersonalized cancer therapyAItheranostics
Y
Yu Gu
Microsoft Research
Yong Xia
Yong Xia
Northwestern Polytechnical University
image processingmedical image analysiscomputer-aided diagnosispattern recognitionmachine learning
Hongming Shan
Hongming Shan
Fudan University; Rensselaer Polytechnic institute
Machine LearningMedical ImagingComputer Vision
Yixuan Yuan
Yixuan Yuan
Associate Professor in Chinese University of Hong Kong
Medical image analysisAI in healthcareBrain data analysisEndoscopy
Yefeng Zheng
Yefeng Zheng
Professor, Westlake University, Hangzhou, China, IEEE Fellow, AIMBE Fellow
AI in HealthMedical ImagingComputer VisionNatural Language ProcessingLarge Language Model
L
Lijun Lu
Southern Medical University