MCSeg: Pre-training and Fine-tuning Volumetric Pyramid Transformer for Multi-modal Cardiac Image Segmentation

๐Ÿ“… 2026-08-31
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บMCSeg๏ผŒไธ€็งๅŸบไบŽไฝ“็งฏๅ˜ๆขๅ™จ็š„็ฝ‘็ปœ๏ผŒ้€š่ฟ‡ๆ–ฐๅž‹็‰นๅพ้‡‘ๅญ—ๅก”ๅ’Œ่‡ช็›‘็ฃ้ข„่ฎญ็ปƒ่งฃๅ†ณๅคšๆจกๆ€ๅฟƒ่„ๅ›พๅƒๅˆ†ๅ‰ฒ้—ฎ้ข˜ใ€‚
๐Ÿ“ Abstract
Automatic cardiac image segmentation is pivotal for diagnosing and treating cardiac diseases. In this work, we introduce MCSeg, a volumetric transformer-based network tailored for multi-modal cardiac segmentation. To overcome the architectural mismatch inherent in existing hybrid networks, we propose a novel Scaling Feature Pyramid (SFP). Unlike conventional skip connections, the SFP effectively bridges the single-scale 3D Vision Transformer (ViT) encoder and the multi-scale CNN decoder by transforming the ViT's output into a hierarchical feature pyramid, ensuring that global contextual information is effectively leveraged. For the training paradigm, the ViT encoder first undergoes self-supervised pre-training via masked image modeling. Subsequently, the network is fine-tuned on downstream tasks, during which a regional mutual information (RMI) loss is integrated to improve boundary segmentation accuracy. In experiments, MCSeg consistently outperforms eleven SOTA methods on CT dataset ImageCHD, multi-modal dataset MM-WHS, MRI dataset HVSMR-2.0 and MSD Heart, highlighting the effectiveness of our MCSeg for multi-modal cardiac segmentation tasks. Furthermore, MCSeg's superior performance in few-shot experiment showcases its significant potential in adapting to limited data scenarios. Codes and pre-trained ViT-B weights are open-sourced at https://openi.pcl.ac.cn/OpenMedIA/MCSeg
Problem

Research questions and friction points this paper is trying to address.

cardiac image segmentation
architectural mismatch
multi-modal
Innovation

Methods, ideas, or system contributions that make the work stand out.

Volumetric Transformer
Scaling Feature Pyramid
Self-supervised Pre-training
Regional Mutual Information Loss
๐Ÿ”Ž Similar Papers
No similar papers found.
Z
Zhiyu Ye
Shenzhen Institute of Advanced Technology, Shenzhen, China; Pengcheng Laboratory, Shenzhen, China; University of Chinese Academy of Sciences, China
Hairong Zheng
Hairong Zheng
Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences
biomedical imaging
T
Tong Zhang
Pengcheng Laboratory, Shenzhen, China