Institution profile

Foshan University

Academic institutionasia · cn
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

IRPol-Fuse: Energy-structure coordination for infrared polarization fusion under low visibility

Aug 07, 2026

This work addresses the limitation of existing infrared-polarization image fusion methods, which overly rely on infrared information under low-visibility conditions, thereby sacrificing polarization texture details. To overcome this, the authors propose IRPol-Fuse, a novel framework that employs an energy-structure collaborative strategy to jointly preserve infrared thermal saliency and polarization structural details. The framework integrates three key components: polarization attention fusion, infrared highlight injection, and polarization texture injection. Additionally, the study introduces LI-PI, the first infrared-polarization benchmark dataset tailored for low-visibility concealed scenarios. Extensive experiments demonstrate that IRPol-Fuse significantly enhances thermal target preservation, structural detail recovery, and visual naturalness on both the LI-PI and LDDRS datasets. The effectiveness of the fused images is further corroborated by improved performance in downstream object detection tasks.

0 citationsRead paper

Degradation-Aware Blur-Segmentation of Brain Tumor

May 15, 2026

This work addresses the challenge of multimodal 3D MRI brain tumor segmentation, which is highly susceptible to motion-induced blurring and artifacts that degrade boundary delineation and segmentation performance. To tackle this issue, the authors propose DABSeg, a novel network that unifies degradation-aware joint deblurring and segmentation for the first time. The method incorporates a feature-domain motion deblurring module to compensate for blur and rebalance intensity, enhanced by blur-aware cross-modal attention and multi-scale residual aggregation to improve feature robustness under degradation. Additionally, it employs a clear-reference reconstruction loss combined with a small-lesion-weighted Dice loss for joint optimization. Evaluated on the BraTS2020 dataset, DABSeg significantly outperforms state-of-the-art methods under both pristine and degraded conditions, achieving notable gains in segmentation accuracy for small lesions and boundary regions.

0 citationsRead paper

FTPFusion: Frequency-Aware Infrared and Visible Video Fusion with Temporal Perturbation

Apr 02, 2026

This work addresses the challenge of simultaneously preserving spatiotemporal consistency and high-frequency details in infrared and visible video fusion. To this end, the authors propose a frequency-aware fusion framework that decomposes features into high- and low-frequency components for separate modeling. The high-frequency branch captures motion cues and fine details through sparse cross-modal spatiotemporal interactions, while the low-frequency branch enhances robustness to dynamic artifacts such as flickering and jitter via a temporal perturbation strategy. Additionally, an offset-aware temporal consistency constraint is introduced to stabilize inter-frame representations. By jointly integrating frequency decomposition, sparse cross-modal interaction, and temporal perturbation—a combination not previously explored—this method achieves state-of-the-art performance on multiple public benchmarks, significantly improving temporal stability without compromising high-frequency detail preservation.

0 citationsRead paper

Text-Guided Channel Perturbation and Pretrained Knowledge Integration for Unified Multi-Modality Image Fusion

Nov 15, 2025

Multimodal image fusion faces two key challenges: gradient conflicts arising from cross-modal parameter sharing, which degrade performance; and modality-specific encoders that improve fusion quality yet harm task generalization. To address these, we propose a unified fusion framework integrating three novel components: semantic-aware channel pruning (to retain discriminative features), geometric affine modulation (to model inter-modal spatial discrepancies), and text-guided channel perturbation (to inject semantic priors and enhance robustness). Our method synergistically leverages pretrained semantic knowledge, channel-level perturbations, and affine transformations—achieving selective feature learning and strong cross-task generalization without introducing modality-specific parameters. Extensive experiments demonstrate consistent and significant improvements over state-of-the-art methods on major fusion benchmarks and downstream detection and segmentation tasks.

0 citationsRead paper

sketch2symm: Symmetry-aware sketch-to-shape generation via semantic bridging

Oct 13, 2025

Sketch-based 3D reconstruction faces significant challenges due to the abstract, sparse nature of input sketches and their limited semantic and geometric information. To address this, we propose a two-stage generative framework that jointly leverages semantic bridging and symmetry constraints. First, a sketch-to-image translation module establishes cross-modal semantic alignment, mitigating information scarcity. Second, a symmetry-aware 3D reconstruction network incorporates reflection symmetry as a strong geometric prior to enhance structural plausibility and completeness. Evaluated on mainstream sketch datasets including ShapeNet, our method achieves state-of-the-art performance across all three core metrics—Chamfer Distance, Earth Mover’s Distance, and F-Score—demonstrating the effectiveness of co-modeling semantic guidance and geometric priors for sketch-based 3D reconstruction.

0 citationsRead paper
Recent publications

Latest Papers

IRPol-Fuse: Energy-structure coordination for infrared polarization fusion under low visibility

Aug 07, 2026

This work addresses the limitation of existing infrared-polarization image fusion methods, which overly rely on infrared information under low-visibility conditions, thereby sacrificing polarization texture details. To overcome this, the authors propose IRPol-Fuse, a novel framework that employs an energy-structure collaborative strategy to jointly preserve infrared thermal saliency and polarization structural details. The framework integrates three key components: polarization attention fusion, infrared highlight injection, and polarization texture injection. Additionally, the study introduces LI-PI, the first infrared-polarization benchmark dataset tailored for low-visibility concealed scenarios. Extensive experiments demonstrate that IRPol-Fuse significantly enhances thermal target preservation, structural detail recovery, and visual naturalness on both the LI-PI and LDDRS datasets. The effectiveness of the fused images is further corroborated by improved performance in downstream object detection tasks.

0 citationsRead paper

Degradation-Aware Blur-Segmentation of Brain Tumor

May 15, 2026

This work addresses the challenge of multimodal 3D MRI brain tumor segmentation, which is highly susceptible to motion-induced blurring and artifacts that degrade boundary delineation and segmentation performance. To tackle this issue, the authors propose DABSeg, a novel network that unifies degradation-aware joint deblurring and segmentation for the first time. The method incorporates a feature-domain motion deblurring module to compensate for blur and rebalance intensity, enhanced by blur-aware cross-modal attention and multi-scale residual aggregation to improve feature robustness under degradation. Additionally, it employs a clear-reference reconstruction loss combined with a small-lesion-weighted Dice loss for joint optimization. Evaluated on the BraTS2020 dataset, DABSeg significantly outperforms state-of-the-art methods under both pristine and degraded conditions, achieving notable gains in segmentation accuracy for small lesions and boundary regions.

0 citationsRead paper

FTPFusion: Frequency-Aware Infrared and Visible Video Fusion with Temporal Perturbation

Apr 02, 2026

This work addresses the challenge of simultaneously preserving spatiotemporal consistency and high-frequency details in infrared and visible video fusion. To this end, the authors propose a frequency-aware fusion framework that decomposes features into high- and low-frequency components for separate modeling. The high-frequency branch captures motion cues and fine details through sparse cross-modal spatiotemporal interactions, while the low-frequency branch enhances robustness to dynamic artifacts such as flickering and jitter via a temporal perturbation strategy. Additionally, an offset-aware temporal consistency constraint is introduced to stabilize inter-frame representations. By jointly integrating frequency decomposition, sparse cross-modal interaction, and temporal perturbation—a combination not previously explored—this method achieves state-of-the-art performance on multiple public benchmarks, significantly improving temporal stability without compromising high-frequency detail preservation.

0 citationsRead paper

Text-Guided Channel Perturbation and Pretrained Knowledge Integration for Unified Multi-Modality Image Fusion

Nov 15, 2025

Multimodal image fusion faces two key challenges: gradient conflicts arising from cross-modal parameter sharing, which degrade performance; and modality-specific encoders that improve fusion quality yet harm task generalization. To address these, we propose a unified fusion framework integrating three novel components: semantic-aware channel pruning (to retain discriminative features), geometric affine modulation (to model inter-modal spatial discrepancies), and text-guided channel perturbation (to inject semantic priors and enhance robustness). Our method synergistically leverages pretrained semantic knowledge, channel-level perturbations, and affine transformations—achieving selective feature learning and strong cross-task generalization without introducing modality-specific parameters. Extensive experiments demonstrate consistent and significant improvements over state-of-the-art methods on major fusion benchmarks and downstream detection and segmentation tasks.

0 citationsRead paper

sketch2symm: Symmetry-aware sketch-to-shape generation via semantic bridging

Oct 13, 2025

Sketch-based 3D reconstruction faces significant challenges due to the abstract, sparse nature of input sketches and their limited semantic and geometric information. To address this, we propose a two-stage generative framework that jointly leverages semantic bridging and symmetry constraints. First, a sketch-to-image translation module establishes cross-modal semantic alignment, mitigating information scarcity. Second, a symmetry-aware 3D reconstruction network incorporates reflection symmetry as a strong geometric prior to enhance structural plausibility and completeness. Evaluated on mainstream sketch datasets including ShapeNet, our method achieves state-of-the-art performance across all three core metrics—Chamfer Distance, Earth Mover’s Distance, and F-Score—demonstrating the effectiveness of co-modeling semantic guidance and geometric priors for sketch-based 3D reconstruction.

0 citationsRead paper