MSA2-Net: Utilizing Self-Adaptive Convolution Module to Extract Multi-Scale Information in Medical Image Segmentation
While nnUNet automatically optimizes training hyperparameters, its fixed internal architecture—particularly static convolutional kernel sizes—limits its capacity to model multi-organ, multi-scale anatomical features. To address this, we propose MSA2-Net: an encoder-decoder framework integrating adaptive convolution modules that dynamically adjust kernel sizes to accommodate organ-scale variations; incorporating CSWin Transformer for long-range dependency modeling; and introducing multi-scale convolutional bridges with optimized skip connections to synergistically enhance global-local feature interaction. Evaluated on Synapse, ACDC, Kvasir, and ISIC2017, MSA2-Net achieves Dice scores of 86.49%, 92.56%, 93.37%, and 92.98%, respectively, demonstrating substantial improvements in generalization and segmentation accuracy. The core contributions are a structural adaptive convolution mechanism and a novel multi-scale Transformer fusion architecture.