Institution profile

Raidium

Industry research
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Curia-MAE: Multi-Modal Multi-Anatomy MAE Pre-Training for 3D Medical Image Segmentation

Aug 06, 2026

This work addresses the limited generalization of existing medical foundation models in 3D dense prediction tasks, where frozen encoders significantly underperform fully trained nnU-Net. To overcome this, the authors propose an enhanced convolutional Masked Autoencoder (MAE) pretraining framework incorporating robust reconstruction targets, feature regularization, and local-global similarity-based contrastive learning. For the first time, this approach enables unified pretraining across large-scale, multimodal (CT/MRI) datasets spanning diverse anatomical regions. Evaluated on eight segmentation benchmarks, the method substantially outperforms strong MAE baselines: with a frozen encoder, it achieves markedly improved performance, while remaining competitive under full fine-tuning—particularly excelling in lesion segmentation tasks with scarce annotations. These results demonstrate its potential to support efficient clinical deployment.

0 citationsRead paper

RAPS-3D: Efficient interactive segmentation for 3D radiological imaging

Jul 10, 2025

Existing SAM-based methods operate on 2D images and cannot directly process 3D medical volumes (e.g., CT/MRI); moreover, autoregressive slice-wise inference and sliding-window strategies incur high computational overhead, latency, and implementation complexity. To address these limitations, we propose an end-to-end, lightweight 3D prompt-driven segmentation framework built upon the SegVol architecture. Our method introduces a compact 3D prompt encoder and a voxel-wise fully convolutional decoder, enabling direct support for interactive prompts—including points and bounding boxes—without sliding windows or sequential modeling. Evaluated across multiple 3D medical imaging benchmarks, our approach achieves state-of-the-art segmentation accuracy while accelerating inference by 2.1–3.8× and reducing GPU memory consumption by 47%–63%. These improvements significantly enhance real-time performance and clinical usability for interactive 3D medical image segmentation.

0 citationsRead paper

RadSAM: Segmenting 3D radiological images with a 2D promptable model

Apr 29, 2025

To address the limitations of slice-wise prompting and lack of interactive editing in 3D medical image segmentation, this paper proposes the first single-prompt-driven framework for 3D segmentation using a 2D foundation model. Methodologically, it introduces (1) noisy masks as a novel weakly supervised prompt type; (2) a slice-wise iterative inference mechanism jointly optimized with 3D consistency constraints; and (3) the first benchmark supporting cross-domain generalization and real-time interactive editing evaluation for single-prompt 3D medical segmentation. Built upon the SAM architecture, the framework fuses multimodal prompts—including noisy masks, points, and bounding boxes—to achieve high-precision 3D organ segmentation from a single prompt on the AMOS dataset, attaining an mDice of 82.7%, substantially outperforming existing methods. It further demonstrates strong cross-domain transferability and clinically viable interactive editing capabilities.

0 citationsRead paper
Recent publications

Latest Papers

Curia-MAE: Multi-Modal Multi-Anatomy MAE Pre-Training for 3D Medical Image Segmentation

Aug 06, 2026

This work addresses the limited generalization of existing medical foundation models in 3D dense prediction tasks, where frozen encoders significantly underperform fully trained nnU-Net. To overcome this, the authors propose an enhanced convolutional Masked Autoencoder (MAE) pretraining framework incorporating robust reconstruction targets, feature regularization, and local-global similarity-based contrastive learning. For the first time, this approach enables unified pretraining across large-scale, multimodal (CT/MRI) datasets spanning diverse anatomical regions. Evaluated on eight segmentation benchmarks, the method substantially outperforms strong MAE baselines: with a frozen encoder, it achieves markedly improved performance, while remaining competitive under full fine-tuning—particularly excelling in lesion segmentation tasks with scarce annotations. These results demonstrate its potential to support efficient clinical deployment.

0 citationsRead paper

RAPS-3D: Efficient interactive segmentation for 3D radiological imaging

Jul 10, 2025

Existing SAM-based methods operate on 2D images and cannot directly process 3D medical volumes (e.g., CT/MRI); moreover, autoregressive slice-wise inference and sliding-window strategies incur high computational overhead, latency, and implementation complexity. To address these limitations, we propose an end-to-end, lightweight 3D prompt-driven segmentation framework built upon the SegVol architecture. Our method introduces a compact 3D prompt encoder and a voxel-wise fully convolutional decoder, enabling direct support for interactive prompts—including points and bounding boxes—without sliding windows or sequential modeling. Evaluated across multiple 3D medical imaging benchmarks, our approach achieves state-of-the-art segmentation accuracy while accelerating inference by 2.1–3.8× and reducing GPU memory consumption by 47%–63%. These improvements significantly enhance real-time performance and clinical usability for interactive 3D medical image segmentation.

0 citationsRead paper

RadSAM: Segmenting 3D radiological images with a 2D promptable model

Apr 29, 2025

To address the limitations of slice-wise prompting and lack of interactive editing in 3D medical image segmentation, this paper proposes the first single-prompt-driven framework for 3D segmentation using a 2D foundation model. Methodologically, it introduces (1) noisy masks as a novel weakly supervised prompt type; (2) a slice-wise iterative inference mechanism jointly optimized with 3D consistency constraints; and (3) the first benchmark supporting cross-domain generalization and real-time interactive editing evaluation for single-prompt 3D medical segmentation. Built upon the SAM architecture, the framework fuses multimodal prompts—including noisy masks, points, and bounding boxes—to achieve high-precision 3D organ segmentation from a single prompt on the AMOS dataset, attaining an mDice of 82.7%, substantially outperforming existing methods. It further demonstrates strong cross-domain transferability and clinically viable interactive editing capabilities.

0 citationsRead paper