Institution profile

Tiangong University

Academic institutionasia · cn
Official website
Research library56linked papers
Opportunities0open roles
Selected work

Representative Papers

CSG-Mamba: A Convolutional Scoring Gating Vision State Space Network for Endoscopic Polyp Segmentation

Aug 14, 2026

This study addresses the challenges of ambiguous boundaries and weak local geometric continuity in visual state space models for endoscopic polyp segmentation by proposing CSG-Mamba. Specifically, a Convolutional Scoring Gating module is embedded within the VM-UNet bottleneck, leveraging large-kernel depthwise and pointwise convolutions to generate spatial score maps. Through multiplicative gating, this mechanism recalibrates features to enhance local modeling capabilities and mitigate contour over-smoothing. Experimental evaluations on the Kvasir-SEG and CVC-ColonDB datasets demonstrate that CSG-Mamba achieves a Dice score of 0.9220 and an mIoU of 0.7418, respectively. These results indicate significant improvements over existing baselines, particularly in boundary precision, validating the effectiveness of integrating convolutional gating into state space architectures for precise medical image segmentation.

0 citationsRead paper

SRE-FER: Regional residual evidence learning for mitigating local evidence dilution in fine-grained facial expression recognition

Aug 09, 2026

This work addresses the challenge in fine-grained facial expression recognition where global feature aggregation often dilutes subtle local muscle cues, leading to confusion between adjacent emotion categories such as fear/surprise and sadness/neutral. To mitigate this, the authors propose the SRE-FER framework, which introduces region-wise residual evidence learning at the readout layer. By employing zero-initialized residual logits, the method sharpens class boundaries while preserving the backbone’s global predictions. A Region-Enhanced Residual Attention (RERA) module leverages Facial Action Coding System (FACS) anatomical priors to guide regional feature focus toward expression-relevant areas, eliminating the need for external keypoints during inference. An optional Full setting dynamically selects non-redundant tokens for efficiency. Built upon a DINOv3 backbone, the approach achieves state-of-the-art accuracy of 92.76%, 91.32%, and 67.78% on RAF-DB, FERPlus, and AffectNet-7, respectively.

0 citationsRead paper

C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs

Jul 23, 2026

This work addresses the significant performance degradation of multimodal large language models under post-training quantization, which stems from outlier channels highly sensitive to quantization. The authors propose a unified channel-level quantization method that, for the first time, incorporates task-specific Fisher information into sensitivity assessment. By introducing a Fisher-weighted objective that jointly accounts for task loss perturbations and quantization error, the method guides channel scaling to preserve critical channels—without requiring auxiliary modules such as LoRA. Combining Hessian approximation, channel-wise scaling, and post-training quantization, the approach achieves state-of-the-art results across eight benchmarks on Qwen2.5-VL, InternVL2, and LLaVA-OV, under both weight-only and weight-activation quantization settings.

0 citationsRead paper

Histopathological Spectrum-Guided Prostate Stratification via Segmentation-Assisted Diagnostic Transformer

Jul 19, 2026

Current MRI-based diagnosis of prostate cancer relies on subjective PI-RADS scoring or coarse binary classification, which fails to capture pathological heterogeneity and is hindered by the scarcity of benign samples. To address these limitations, this work introduces PCa-HSD, the first histopathological spectrum dataset for prostate cancer, and formulates a fine-grained four-class risk stratification task. The authors propose a Language-guided Segmentation-assisted Diagnostic Transformer (LSDT) that incorporates anatomical priors via zero-shot segmentation and fuses multimodal MRI slices. Evaluated on 344 patients using five-fold cross-validation, the model achieves an average accuracy of 0.633 and a JointRecall of 0.768, significantly outperforming baseline methods. This approach effectively mitigates the challenge of insufficient benign samples and enhances clinical relevance.

0 citationsRead paper
Recent publications

Latest Papers

CSG-Mamba: A Convolutional Scoring Gating Vision State Space Network for Endoscopic Polyp Segmentation

Aug 14, 2026

This study addresses the challenges of ambiguous boundaries and weak local geometric continuity in visual state space models for endoscopic polyp segmentation by proposing CSG-Mamba. Specifically, a Convolutional Scoring Gating module is embedded within the VM-UNet bottleneck, leveraging large-kernel depthwise and pointwise convolutions to generate spatial score maps. Through multiplicative gating, this mechanism recalibrates features to enhance local modeling capabilities and mitigate contour over-smoothing. Experimental evaluations on the Kvasir-SEG and CVC-ColonDB datasets demonstrate that CSG-Mamba achieves a Dice score of 0.9220 and an mIoU of 0.7418, respectively. These results indicate significant improvements over existing baselines, particularly in boundary precision, validating the effectiveness of integrating convolutional gating into state space architectures for precise medical image segmentation.

0 citationsRead paper

SRE-FER: Regional residual evidence learning for mitigating local evidence dilution in fine-grained facial expression recognition

Aug 09, 2026

This work addresses the challenge in fine-grained facial expression recognition where global feature aggregation often dilutes subtle local muscle cues, leading to confusion between adjacent emotion categories such as fear/surprise and sadness/neutral. To mitigate this, the authors propose the SRE-FER framework, which introduces region-wise residual evidence learning at the readout layer. By employing zero-initialized residual logits, the method sharpens class boundaries while preserving the backbone’s global predictions. A Region-Enhanced Residual Attention (RERA) module leverages Facial Action Coding System (FACS) anatomical priors to guide regional feature focus toward expression-relevant areas, eliminating the need for external keypoints during inference. An optional Full setting dynamically selects non-redundant tokens for efficiency. Built upon a DINOv3 backbone, the approach achieves state-of-the-art accuracy of 92.76%, 91.32%, and 67.78% on RAF-DB, FERPlus, and AffectNet-7, respectively.

0 citationsRead paper

C-PTQ: Fisher-weighted Channel-wise Sensitivity for Post-training Quantization of MLLMs

Jul 23, 2026

This work addresses the significant performance degradation of multimodal large language models under post-training quantization, which stems from outlier channels highly sensitive to quantization. The authors propose a unified channel-level quantization method that, for the first time, incorporates task-specific Fisher information into sensitivity assessment. By introducing a Fisher-weighted objective that jointly accounts for task loss perturbations and quantization error, the method guides channel scaling to preserve critical channels—without requiring auxiliary modules such as LoRA. Combining Hessian approximation, channel-wise scaling, and post-training quantization, the approach achieves state-of-the-art results across eight benchmarks on Qwen2.5-VL, InternVL2, and LLaVA-OV, under both weight-only and weight-activation quantization settings.

0 citationsRead paper

Histopathological Spectrum-Guided Prostate Stratification via Segmentation-Assisted Diagnostic Transformer

Jul 19, 2026

Current MRI-based diagnosis of prostate cancer relies on subjective PI-RADS scoring or coarse binary classification, which fails to capture pathological heterogeneity and is hindered by the scarcity of benign samples. To address these limitations, this work introduces PCa-HSD, the first histopathological spectrum dataset for prostate cancer, and formulates a fine-grained four-class risk stratification task. The authors propose a Language-guided Segmentation-assisted Diagnostic Transformer (LSDT) that incorporates anatomical priors via zero-shot segmentation and fuses multimodal MRI slices. Evaluated on 344 patients using five-fold cross-validation, the model achieves an average accuracy of 0.633 and a JointRecall of 0.768, significantly outperforming baseline methods. This approach effectively mitigates the challenge of insufficient benign samples and enhances clinical relevance.

0 citationsRead paper