MSCA-UNet: Multi-Scale Context and Attention U-Net for Image Segmentation
本文提出MSCA-UNet,通过结合多尺度上下文聚合和通道-空间注意力机制改进U-Net图像分割网络,有效提升分割精度。
本文提出MSCA-UNet,通过结合多尺度上下文聚合和通道-空间注意力机制改进U-Net图像分割网络,有效提升分割精度。
该研究提出SFMformer,通过在注意力机制前后分别增强空间和频域信息来提高轻量级图像超分辨率质量。
This work addresses the high computational cost of the C2PSA module in the lightweight, NMS-free object detector YOLO26 by introducing the state space model Mamba into its architecture for the first time. Specifically, C2PSA is replaced with MambaPSA at the end of the backbone, and bidirectional Vision Mamba (BiViM) modules are embedded into the P3–P5 layers of the neck. This design achieves a superior trade-off between efficiency and accuracy, reducing parameters by 2.9% and FLOPs by 12.1% on PASCAL VOC, while accelerating CPU inference by 17.6% (from 17 to 20 FPS) with only a marginal 0.1 drop in mAP50:95. Notably, the BiViM module at the P4 layer alone contributes a +0.9 gain in mAP50:95, demonstrating the effectiveness and efficiency of Mamba within NMS-free detection frameworks.
This work addresses the limited cross-domain robustness of document image binarization under unknown degradation conditions such as paper aging, ink bleed-through, stains, shadows, and non-uniform illumination. To this end, we propose DR-Mamba, a novel framework that reconfigures Mamba into a dual-path slow-fast architecture. It explicitly disentangles foreground details from background content via an input-dependent subtraction gating mechanism, enabling inference-time adaptation without requiring target-domain labels or parameter updates within a single forward pass. Coupled with full-resolution detail-guided reconstruction and fine-stroke-aware supervision, DR-Mamba achieves substantial improvements in cross-domain performance under the DIBCO leave-one-out protocol, particularly excelling in scenarios involving the most severe degradations.
This work addresses the degradation of boundary and fine-detail responses in Mamba-based models for multi-class semantic segmentation, which arises from sequential state propagation and leads to reduced segmentation accuracy. To mitigate this issue, the authors propose Reload-Mamba, a novel framework that introduces an anti-dilution mechanism into dense prediction tasks for the first time. It restores detailed representations through a top-down hierarchy across three decoder stages, leveraging boundary-supervised local detail priors, class-uncertainty-aware reloading gates, and a multi-level hierarchical reloading structure. Integrated with a ConvNeXt-Tiny encoder, four-directional Mamba scanning, and pixel-wise directional attention, the model achieves state-of-the-art performance on ADE20K (48.9% mIoU), Cityscapes (83.2%), and PASCAL VOC 2012 (87.8%, a +2.2% improvement), demonstrating its effectiveness in recovering boundary fidelity and fine-grained details.
本文提出MSCA-UNet,通过结合多尺度上下文聚合和通道-空间注意力机制改进U-Net图像分割网络,有效提升分割精度。
该研究提出SFMformer,通过在注意力机制前后分别增强空间和频域信息来提高轻量级图像超分辨率质量。
This work addresses the high computational cost of the C2PSA module in the lightweight, NMS-free object detector YOLO26 by introducing the state space model Mamba into its architecture for the first time. Specifically, C2PSA is replaced with MambaPSA at the end of the backbone, and bidirectional Vision Mamba (BiViM) modules are embedded into the P3–P5 layers of the neck. This design achieves a superior trade-off between efficiency and accuracy, reducing parameters by 2.9% and FLOPs by 12.1% on PASCAL VOC, while accelerating CPU inference by 17.6% (from 17 to 20 FPS) with only a marginal 0.1 drop in mAP50:95. Notably, the BiViM module at the P4 layer alone contributes a +0.9 gain in mAP50:95, demonstrating the effectiveness and efficiency of Mamba within NMS-free detection frameworks.
This work addresses the limited cross-domain robustness of document image binarization under unknown degradation conditions such as paper aging, ink bleed-through, stains, shadows, and non-uniform illumination. To this end, we propose DR-Mamba, a novel framework that reconfigures Mamba into a dual-path slow-fast architecture. It explicitly disentangles foreground details from background content via an input-dependent subtraction gating mechanism, enabling inference-time adaptation without requiring target-domain labels or parameter updates within a single forward pass. Coupled with full-resolution detail-guided reconstruction and fine-stroke-aware supervision, DR-Mamba achieves substantial improvements in cross-domain performance under the DIBCO leave-one-out protocol, particularly excelling in scenarios involving the most severe degradations.
This work addresses the degradation of boundary and fine-detail responses in Mamba-based models for multi-class semantic segmentation, which arises from sequential state propagation and leads to reduced segmentation accuracy. To mitigate this issue, the authors propose Reload-Mamba, a novel framework that introduces an anti-dilution mechanism into dense prediction tasks for the first time. It restores detailed representations through a top-down hierarchy across three decoder stages, leveraging boundary-supervised local detail priors, class-uncertainty-aware reloading gates, and a multi-level hierarchical reloading structure. Integrated with a ConvNeXt-Tiny encoder, four-directional Mamba scanning, and pixel-wise directional attention, the model achieves state-of-the-art performance on ADE20K (48.9% mIoU), Cityscapes (83.2%), and PASCAL VOC 2012 (87.8%, a +2.2% improvement), demonstrating its effectiveness in recovering boundary fidelity and fine-grained details.