Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging
Current AI models struggle to support natural language interaction, interpretability, and cross-modal spatial reasoning in multiparametric 3D MRI, limiting their performance in tasks such as glioma grading. This work proposes the first general-purpose vision–language foundation model tailored for multiparametric 3D MRI, which jointly models modality and spatial information through a shared 3D encoder, 4D rotary positional embeddings, and a multi-resolution feature injection mechanism to enable cross-scale perception. Pretrained unsupervised with 4 billion parameters, the model significantly outperforms existing general and domain-specific large models on report generation (BERTScore: 0.856), visual question answering (accuracy: 0.713), and multiple-choice tasks (accuracy: 0.912).