A Momentum-Based Variance-Reduced Algorithm for Federated Multiobjective Optimization
本文针对联邦多目标优化问题,提出了一种基于动量的方差减少算法,通过改进梯度估计器减少了随机更新的方差,提高了收敛速度。
本文针对联邦多目标优化问题,提出了一种基于动量的方差减少算法,通过改进梯度估计器减少了随机更新的方差,提高了收敛速度。
This work addresses the challenges of channel redundancy and high computational cost in RGB-infrared multimodal object detection arising from parallel feature extraction, where existing pruning methods overlook cross-modal interactions and scene-level dynamic redundancy. To this end, we propose the first interactive structured channel pruning framework tailored for this task, which innovatively integrates three key components: a Taylor-based implicit criterion to quantify channel importance, a Modality Interaction Redundancy Analysis (MIRA) module to model cross-modal complementarity, and a language prior-guided Scene-level Pruning with Contextual Awareness (SPCA) mechanism to enable dynamic, context-aware channel pruning. Evaluated on the FLIR dataset, our method achieves a 0.6% increase in mAP after pruning 50% of channels, demonstrating simultaneous reductions in computational overhead and improvements—or at least preservation—of detection performance.
This work addresses the challenges of ship detection in synthetic aperture radar (SAR) imagery, which is hindered by speckle noise, complex coastal clutter, and the prevalence of small targets—factors that compromise the robustness and fine-grained feature preservation of conventional optical detectors. To overcome these limitations, the authors propose a domain-aware detection model built upon the DETR framework. The approach introduces two key innovations: a SARESMoE module that employs a sparse gating mechanism to dynamically route features between frequency-domain and wavelet-domain experts, and a Space-to-Depth Enhanced Pyramid (SDEP) that effectively integrates high-resolution shallow spatial features to improve small-target localization. Evaluated on benchmarks including HRSID, the model substantially outperforms YOLO variants and existing SAR-specific detectors, achieving an mAP50:95 of 76.4% and an mAP50 of 93.8%.
Existing training-free open-vocabulary remote sensing segmentation methods suffer from inconsistent predictions and limited generalization due to their reliance on independent inference strategies that neglect strong spatial and semantic correlations within images. To address this, this work proposes the first training-free, context-aware reasoning framework that explicitly models semantic dependencies through a joint inference mechanism across spatial regions and integrates global contextual information to enhance segmentation consistency and robustness. Built upon vision-language models, the proposed method significantly outperforms current state-of-the-art approaches across multiple benchmarks, achieving average performance gains of 2.80% in open-vocabulary semantic segmentation and 6.13% in object extraction tasks.
To address severe image degradation—including color distortion, low contrast, and blur—in complex underwater scenes, this paper proposes a multi-scale feature collaborative enhancement model integrating VGG19 and ResNet50. The method jointly optimizes structural detail preservation and accurate color restoration via complementary deep feature extraction within a unified end-to-end training framework. A novel cross-network feature fusion mechanism is introduced to synergistically leverage hierarchical representations from both backbones. Comprehensive quantitative evaluation is conducted using multiple objective metrics—PSNR, UCIQE, and UIQM—on diverse real-world underwater datasets. Experimental results demonstrate that the proposed approach significantly outperforms conventional enhancement methods and single-backbone baselines: it achieves an average UIQM improvement of 12.7%, exhibits strong robustness across heterogeneous underwater conditions, and maintains computational efficiency suitable for practical deployment.
本文针对联邦多目标优化问题,提出了一种基于动量的方差减少算法,通过改进梯度估计器减少了随机更新的方差,提高了收敛速度。
This work addresses the challenges of channel redundancy and high computational cost in RGB-infrared multimodal object detection arising from parallel feature extraction, where existing pruning methods overlook cross-modal interactions and scene-level dynamic redundancy. To this end, we propose the first interactive structured channel pruning framework tailored for this task, which innovatively integrates three key components: a Taylor-based implicit criterion to quantify channel importance, a Modality Interaction Redundancy Analysis (MIRA) module to model cross-modal complementarity, and a language prior-guided Scene-level Pruning with Contextual Awareness (SPCA) mechanism to enable dynamic, context-aware channel pruning. Evaluated on the FLIR dataset, our method achieves a 0.6% increase in mAP after pruning 50% of channels, demonstrating simultaneous reductions in computational overhead and improvements—or at least preservation—of detection performance.
This work addresses the challenges of ship detection in synthetic aperture radar (SAR) imagery, which is hindered by speckle noise, complex coastal clutter, and the prevalence of small targets—factors that compromise the robustness and fine-grained feature preservation of conventional optical detectors. To overcome these limitations, the authors propose a domain-aware detection model built upon the DETR framework. The approach introduces two key innovations: a SARESMoE module that employs a sparse gating mechanism to dynamically route features between frequency-domain and wavelet-domain experts, and a Space-to-Depth Enhanced Pyramid (SDEP) that effectively integrates high-resolution shallow spatial features to improve small-target localization. Evaluated on benchmarks including HRSID, the model substantially outperforms YOLO variants and existing SAR-specific detectors, achieving an mAP50:95 of 76.4% and an mAP50 of 93.8%.
Existing training-free open-vocabulary remote sensing segmentation methods suffer from inconsistent predictions and limited generalization due to their reliance on independent inference strategies that neglect strong spatial and semantic correlations within images. To address this, this work proposes the first training-free, context-aware reasoning framework that explicitly models semantic dependencies through a joint inference mechanism across spatial regions and integrates global contextual information to enhance segmentation consistency and robustness. Built upon vision-language models, the proposed method significantly outperforms current state-of-the-art approaches across multiple benchmarks, achieving average performance gains of 2.80% in open-vocabulary semantic segmentation and 6.13% in object extraction tasks.
To address severe image degradation—including color distortion, low contrast, and blur—in complex underwater scenes, this paper proposes a multi-scale feature collaborative enhancement model integrating VGG19 and ResNet50. The method jointly optimizes structural detail preservation and accurate color restoration via complementary deep feature extraction within a unified end-to-end training framework. A novel cross-network feature fusion mechanism is introduced to synergistically leverage hierarchical representations from both backbones. Comprehensive quantitative evaluation is conducted using multiple objective metrics—PSNR, UCIQE, and UIQM—on diverse real-world underwater datasets. Experimental results demonstrate that the proposed approach significantly outperforms conventional enhancement methods and single-backbone baselines: it achieves an average UIQM improvement of 12.7%, exhibits strong robustness across heterogeneous underwater conditions, and maintains computational efficiency suitable for practical deployment.