Lesion Segmentation in FDG-PET/CT Using Swin Transformer U-Net 3D: A Robust Deep Learning Framework
This study addresses the limited accuracy of automatic lesion segmentation in FDG-PET/CT imaging by proposing SwinUNet3D, a novel framework that effectively integrates the shifted-window self-attention mechanism of Swin Transformer with the skip-connection architecture of 3D U-Net. This integration enables simultaneous modeling of global contextual information and preservation of fine anatomical details, while also optimizing multimodal PET/CT feature fusion. The method substantially enhances detection of small and irregular lesions and reduces false-positive rates. Evaluated on the AutoPET III dataset, SwinUNet3D achieves a Dice coefficient of 0.88 and an IoU of 0.78—significantly outperforming standard 3D U-Net (Dice 0.48, IoU 0.32)—and demonstrates faster inference speed.