A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging
This study addresses the lack of a standardized benchmark for fair comparison among deep learning models in multi-task 3D brain tumor segmentation. For the first time, it systematically evaluates the performance of five prominent architectures—3D U-Net, SegResNet, Swin UNETR, SegMamba, and SegMambaV2—under identical conditions, including the same dataset (BraTS 2023/2024), preprocessing pipeline, training protocol, and evaluation metrics. Through comprehensive assessment across multiple dimensions—segmentation accuracy, inference time, and model size—the work elucidates the inherent trade-offs between accuracy and efficiency. These findings provide clinically relevant guidance for model selection and deployment, effectively establishing a much-needed standardized reference framework in the field.