Nonparametric Change-Point Detection and Inference for High-Dimensional Distributions
本文提出非参数方法,通过标准化秩比较解决高维分布变化点检测问题,适用于未知稀疏性情况,无需边际矩假设。
本文提出非参数方法,通过标准化秩比较解决高维分布变化点检测问题,适用于未知稀疏性情况,无需边际矩假设。
This study addresses the challenges of ambiguous boundaries and weak local geometric continuity in visual state space models for endoscopic polyp segmentation by proposing CSG-Mamba. Specifically, a Convolutional Scoring Gating module is embedded within the VM-UNet bottleneck, leveraging large-kernel depthwise and pointwise convolutions to generate spatial score maps. Through multiplicative gating, this mechanism recalibrates features to enhance local modeling capabilities and mitigate contour over-smoothing. Experimental evaluations on the Kvasir-SEG and CVC-ColonDB datasets demonstrate that CSG-Mamba achieves a Dice score of 0.9220 and an mIoU of 0.7418, respectively. These results indicate significant improvements over existing baselines, particularly in boundary precision, validating the effectiveness of integrating convolutional gating into state space architectures for precise medical image segmentation.
This work addresses the challenge in fine-grained facial expression recognition where global feature aggregation often dilutes subtle local muscle cues, leading to confusion between adjacent emotion categories such as fear/surprise and sadness/neutral. To mitigate this, the authors propose the SRE-FER framework, which introduces region-wise residual evidence learning at the readout layer. By employing zero-initialized residual logits, the method sharpens class boundaries while preserving the backbone’s global predictions. A Region-Enhanced Residual Attention (RERA) module leverages Facial Action Coding System (FACS) anatomical priors to guide regional feature focus toward expression-relevant areas, eliminating the need for external keypoints during inference. An optional Full setting dynamically selects non-redundant tokens for efficiency. Built upon a DINOv3 backbone, the approach achieves state-of-the-art accuracy of 92.76%, 91.32%, and 67.78% on RAF-DB, FERPlus, and AffectNet-7, respectively.
This work addresses the significant performance degradation of multimodal large language models under post-training quantization, which stems from outlier channels highly sensitive to quantization. The authors propose a unified channel-level quantization method that, for the first time, incorporates task-specific Fisher information into sensitivity assessment. By introducing a Fisher-weighted objective that jointly accounts for task loss perturbations and quantization error, the method guides channel scaling to preserve critical channels—without requiring auxiliary modules such as LoRA. Combining Hessian approximation, channel-wise scaling, and post-training quantization, the approach achieves state-of-the-art results across eight benchmarks on Qwen2.5-VL, InternVL2, and LLaVA-OV, under both weight-only and weight-activation quantization settings.
Current MRI-based diagnosis of prostate cancer relies on subjective PI-RADS scoring or coarse binary classification, which fails to capture pathological heterogeneity and is hindered by the scarcity of benign samples. To address these limitations, this work introduces PCa-HSD, the first histopathological spectrum dataset for prostate cancer, and formulates a fine-grained four-class risk stratification task. The authors propose a Language-guided Segmentation-assisted Diagnostic Transformer (LSDT) that incorporates anatomical priors via zero-shot segmentation and fuses multimodal MRI slices. Evaluated on 344 patients using five-fold cross-validation, the model achieves an average accuracy of 0.633 and a JointRecall of 0.768, significantly outperforming baseline methods. This approach effectively mitigates the challenge of insufficient benign samples and enhances clinical relevance.
本文提出非参数方法,通过标准化秩比较解决高维分布变化点检测问题,适用于未知稀疏性情况,无需边际矩假设。
This study addresses the challenges of ambiguous boundaries and weak local geometric continuity in visual state space models for endoscopic polyp segmentation by proposing CSG-Mamba. Specifically, a Convolutional Scoring Gating module is embedded within the VM-UNet bottleneck, leveraging large-kernel depthwise and pointwise convolutions to generate spatial score maps. Through multiplicative gating, this mechanism recalibrates features to enhance local modeling capabilities and mitigate contour over-smoothing. Experimental evaluations on the Kvasir-SEG and CVC-ColonDB datasets demonstrate that CSG-Mamba achieves a Dice score of 0.9220 and an mIoU of 0.7418, respectively. These results indicate significant improvements over existing baselines, particularly in boundary precision, validating the effectiveness of integrating convolutional gating into state space architectures for precise medical image segmentation.
This work addresses the challenge in fine-grained facial expression recognition where global feature aggregation often dilutes subtle local muscle cues, leading to confusion between adjacent emotion categories such as fear/surprise and sadness/neutral. To mitigate this, the authors propose the SRE-FER framework, which introduces region-wise residual evidence learning at the readout layer. By employing zero-initialized residual logits, the method sharpens class boundaries while preserving the backbone’s global predictions. A Region-Enhanced Residual Attention (RERA) module leverages Facial Action Coding System (FACS) anatomical priors to guide regional feature focus toward expression-relevant areas, eliminating the need for external keypoints during inference. An optional Full setting dynamically selects non-redundant tokens for efficiency. Built upon a DINOv3 backbone, the approach achieves state-of-the-art accuracy of 92.76%, 91.32%, and 67.78% on RAF-DB, FERPlus, and AffectNet-7, respectively.
This work addresses the significant performance degradation of multimodal large language models under post-training quantization, which stems from outlier channels highly sensitive to quantization. The authors propose a unified channel-level quantization method that, for the first time, incorporates task-specific Fisher information into sensitivity assessment. By introducing a Fisher-weighted objective that jointly accounts for task loss perturbations and quantization error, the method guides channel scaling to preserve critical channels—without requiring auxiliary modules such as LoRA. Combining Hessian approximation, channel-wise scaling, and post-training quantization, the approach achieves state-of-the-art results across eight benchmarks on Qwen2.5-VL, InternVL2, and LLaVA-OV, under both weight-only and weight-activation quantization settings.
Current MRI-based diagnosis of prostate cancer relies on subjective PI-RADS scoring or coarse binary classification, which fails to capture pathological heterogeneity and is hindered by the scarcity of benign samples. To address these limitations, this work introduces PCa-HSD, the first histopathological spectrum dataset for prostate cancer, and formulates a fine-grained four-class risk stratification task. The authors propose a Language-guided Segmentation-assisted Diagnostic Transformer (LSDT) that incorporates anatomical priors via zero-shot segmentation and fuses multimodal MRI slices. Evaluated on 344 patients using five-fold cross-validation, the model achieves an average accuracy of 0.633 and a JointRecall of 0.768, significantly outperforming baseline methods. This approach effectively mitigates the challenge of insufficient benign samples and enhances clinical relevance.