Curia-MAE: Multi-Modal Multi-Anatomy MAE Pre-Training for 3D Medical Image Segmentation

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited generalization of existing medical foundation models in 3D dense prediction tasks, where frozen encoders significantly underperform fully trained nnU-Net. To overcome this, the authors propose an enhanced convolutional Masked Autoencoder (MAE) pretraining framework incorporating robust reconstruction targets, feature regularization, and local-global similarity-based contrastive learning. For the first time, this approach enables unified pretraining across large-scale, multimodal (CT/MRI) datasets spanning diverse anatomical regions. Evaluated on eight segmentation benchmarks, the method substantially outperforms strong MAE baselines: with a frozen encoder, it achieves markedly improved performance, while remaining competitive under full fine-tuning—particularly excelling in lesion segmentation tasks with scarce annotations. These results demonstrate its potential to support efficient clinical deployment.
📝 Abstract
Radiology foundation models learn transferable representations that can be adapted to new tasks by training only small layers on top of a frozen encoder. Dense prediction tasks such as 3D segmentation are, however, underrepresented in their evaluation, and, with the encoder kept frozen, pre-trained models still fall short of nnU-Net, the state-of-the-art reference trained from scratch. To close this gap we extend convolutional MAE pre-training with a robust reconstruction objective, a feature regularizer, and a local-global similarity objective. Using this method, we propose Curia-MAE, a multi-modal, multi-anatomy MAE model pre-trained on 300,000 CT and MRI images covering a large number of anatomical sites. On eight anatomy- and lesion-focused segmentation benchmarks, Curia-MAE improves frozen-encoder performance over a strong MAE baseline, while remaining competitive under full finetuning and superior on lesion tasks, where labeled data is scarce. These results indicate that a single frozen encoder can be reused across diverse segmentation tasks, reducing the cost of adapting and deploying such models in clinical workflows. We will make our pre-trained model weights publicly available.
Problem

Research questions and friction points this paper is trying to address.

3D medical image segmentation
foundation models
dense prediction
transfer learning
multi-anatomy
Innovation

Methods, ideas, or system contributions that make the work stand out.

masked autoencoder
multi-modal pre-training
3D medical image segmentation
frozen encoder
foundation model
T
Théo Danielou
Raidium
A
Antoine Saporta
Raidium
L
Léo Alberge
Raidium
Corentin Dancette
Corentin Dancette
Raidium
Deep LearningVisual Question AnsweringBiasesComputer VisionMedical Imaging