Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current AI models struggle to support natural language interaction, interpretability, and cross-modal spatial reasoning in multiparametric 3D MRI, limiting their performance in tasks such as glioma grading. This work proposes the first general-purpose vision–language foundation model tailored for multiparametric 3D MRI, which jointly models modality and spatial information through a shared 3D encoder, 4D rotary positional embeddings, and a multi-resolution feature injection mechanism to enable cross-scale perception. Pretrained unsupervised with 4 billion parameters, the model significantly outperforms existing general and domain-specific large models on report generation (BERTScore: 0.856), visual question answering (accuracy: 0.713), and multiple-choice tasks (accuracy: 0.912).
📝 Abstract
Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face critical limitations: their lack of natural language interaction and interpretability impedes spatial information integration and cross-modal reasoning required clinically. Key challenges arise from significant physical meaning differences across modalities, spatial misalignment due to scan intervals, and the need for complex multi-feature interpretation in tasks like glioma grading. While visual-language models (VLMs) show promise in cross-modal understanding, existing methods focus mainly on 2D image modeling, neglecting direct perception of 3D volumetric space. Although 3D VLMs have been proposed for report generation and feature alignment in 3D CT imaging, mpMRI applications demand collaborative inference across multiple imaging modalities-a requirement unmet by current solutions. To address this, we introduce Mr3D-VL, a dedicated visual-language foundation model for multi-parametric 3D MRI. With 4 billion parameters, it employs an unsupervised pre-trained shared 3D encoder and 4D rotational positional embedding for dual modality-spatial integration. Its cross-modal projection layer uses a multi-resolution feature implantation strategy to enhance feature perception across resolutions. Experimental results show significant improvements over existing 4B/7B/30B domain-specific and general-purpose models in text generation tasks, achieving a BERTScore of 0.856 for report generation, with question-answering accuracy at 0.713 and multiple-choice accuracy at 0.912.
Problem

Research questions and friction points this paper is trying to address.

multiparametric MRI
3D vision-language model
cross-modal reasoning
spatial misalignment
glioma grading
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D visual-language model
multiparametric MRI
4D rotational positional embedding
multi-resolution feature implantation
foundation model
🔎 Similar Papers
Z
Zhi Qiao
United Imaging Intelligence, Shanghai, China
X
Xintong Wu
United Imaging Intelligence, Shanghai, China
Y
Yichu He
United Imaging Intelligence, Shanghai, China
Feng Shi
Feng Shi
United Imaging Intelligence, Shanghai, China