Institution profile

Inner Mongolia University of Technology

Academic institutionasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

ClickRemoval: An Interactive Open-Source Tool for Object Removal in Diffusion Models

May 14, 2026

Existing object removal methods rely on manual masks or textual prompts, which often struggle to achieve precise manipulation in complex scenes, leading to incomplete removal or background distortion. This work proposes a click-based object removal approach that requires only user-provided point clicks—eliminating the need for additional training, hand-drawn masks, or textual descriptions. By leveraging a pre-trained Stable Diffusion model, the method performs target localization and background inpainting directly during the denoising process through self-attention modulation. To the best of our knowledge, this is the first technique to enable purely click-driven image editing with diffusion models, substantially lowering the usability barrier. The approach demonstrates superior performance in both quantitative evaluations and user studies, and the complete software package has been publicly released.

0 citationsRead paper

DB SwinT: A Dual-Branch Swin Transformer Network for Road Extraction in Optical Remote Sensing Imagery

Mar 25, 2026

This work addresses the challenge of fragmented road structures and low extraction accuracy in optical remote sensing imagery caused by occlusions from trees, buildings, and other objects. To this end, the authors propose a dual-branch Swin Transformer network that integrates a U-Net–inspired multi-scale feature fusion strategy. The architecture employs separate local and global branches to recover fine details in occluded regions and preserve topological continuity of road networks, respectively. An Attention-based Feature Fusion (AFF) module is further introduced to adaptively integrate information from both branches. This design effectively balances local detail reconstruction with global semantic context modeling. Experimental results demonstrate state-of-the-art performance, achieving Intersection over Union (IoU) scores of 79.35% and 74.84% on the Massachusetts and DeepGlobe road datasets, respectively, significantly outperforming existing methods.

0 citationsRead paper

iCD: A Implicit Clustering Distillation Mathod for Structural Information Mining

Sep 15, 2025

Logit-based knowledge distillation, while computationally efficient, suffers from poor interpretability and reliance on ground-truth labels or intermediate feature alignment. To address these limitations, we propose Implicit Clustering Distillation (iCD), the first method to uncover implicit clustering structure directly from decoupled local logit representations. iCD models semantic correlations among logits via a Gram matrix, enabling structured knowledge transfer without label supervision or explicit feature-level alignment. This approach enhances both the student model’s capacity to capture latent semantic structures and the interpretability of its decision-making process. Extensive experiments across multiple benchmark datasets demonstrate that iCD consistently outperforms state-of-the-art baselines, achieving up to a 5.08% absolute accuracy gain—particularly pronounced in fine-grained classification tasks where semantic granularity is critical.

0 citationsRead paper

Dual Consistent Constraint via Disentangled Consistency and Complementarity for Multi-view Clustering

Apr 07, 2025

Existing multi-view clustering methods overemphasize representation consistency across views while neglecting view complementarity. Method: This paper proposes a Dual Consistency Constraint framework built upon a disentangled variational autoencoder that explicitly separates shared (consistent) and private (complementary) representations. It jointly optimizes consistency and complementarity via cross-view contrastive learning, intra- and inter-reconstruction of shared/private information, and mutual information maximization. Contribution/Results: To our knowledge, this is the first work to model and jointly exploit both consistency and complementarity under a unified theoretical framework, breaking away from conventional single-consistency paradigms. Extensive experiments on multiple benchmark datasets demonstrate significant improvements over state-of-the-art methods, validating simultaneous enhancement in representation discriminability and generalizability. The framework also exhibits strong scalability.

0 citationsRead paper
Recent publications

Latest Papers

ClickRemoval: An Interactive Open-Source Tool for Object Removal in Diffusion Models

May 14, 2026

Existing object removal methods rely on manual masks or textual prompts, which often struggle to achieve precise manipulation in complex scenes, leading to incomplete removal or background distortion. This work proposes a click-based object removal approach that requires only user-provided point clicks—eliminating the need for additional training, hand-drawn masks, or textual descriptions. By leveraging a pre-trained Stable Diffusion model, the method performs target localization and background inpainting directly during the denoising process through self-attention modulation. To the best of our knowledge, this is the first technique to enable purely click-driven image editing with diffusion models, substantially lowering the usability barrier. The approach demonstrates superior performance in both quantitative evaluations and user studies, and the complete software package has been publicly released.

0 citationsRead paper

DB SwinT: A Dual-Branch Swin Transformer Network for Road Extraction in Optical Remote Sensing Imagery

Mar 25, 2026

This work addresses the challenge of fragmented road structures and low extraction accuracy in optical remote sensing imagery caused by occlusions from trees, buildings, and other objects. To this end, the authors propose a dual-branch Swin Transformer network that integrates a U-Net–inspired multi-scale feature fusion strategy. The architecture employs separate local and global branches to recover fine details in occluded regions and preserve topological continuity of road networks, respectively. An Attention-based Feature Fusion (AFF) module is further introduced to adaptively integrate information from both branches. This design effectively balances local detail reconstruction with global semantic context modeling. Experimental results demonstrate state-of-the-art performance, achieving Intersection over Union (IoU) scores of 79.35% and 74.84% on the Massachusetts and DeepGlobe road datasets, respectively, significantly outperforming existing methods.

0 citationsRead paper

iCD: A Implicit Clustering Distillation Mathod for Structural Information Mining

Sep 15, 2025

Logit-based knowledge distillation, while computationally efficient, suffers from poor interpretability and reliance on ground-truth labels or intermediate feature alignment. To address these limitations, we propose Implicit Clustering Distillation (iCD), the first method to uncover implicit clustering structure directly from decoupled local logit representations. iCD models semantic correlations among logits via a Gram matrix, enabling structured knowledge transfer without label supervision or explicit feature-level alignment. This approach enhances both the student model’s capacity to capture latent semantic structures and the interpretability of its decision-making process. Extensive experiments across multiple benchmark datasets demonstrate that iCD consistently outperforms state-of-the-art baselines, achieving up to a 5.08% absolute accuracy gain—particularly pronounced in fine-grained classification tasks where semantic granularity is critical.

0 citationsRead paper

Dual Consistent Constraint via Disentangled Consistency and Complementarity for Multi-view Clustering

Apr 07, 2025

Existing multi-view clustering methods overemphasize representation consistency across views while neglecting view complementarity. Method: This paper proposes a Dual Consistency Constraint framework built upon a disentangled variational autoencoder that explicitly separates shared (consistent) and private (complementary) representations. It jointly optimizes consistency and complementarity via cross-view contrastive learning, intra- and inter-reconstruction of shared/private information, and mutual information maximization. Contribution/Results: To our knowledge, this is the first work to model and jointly exploit both consistency and complementarity under a unified theoretical framework, breaking away from conventional single-consistency paradigms. Extensive experiments on multiple benchmark datasets demonstrate significant improvements over state-of-the-art methods, validating simultaneous enhancement in representation discriminability and generalizability. The framework also exhibits strong scalability.

0 citationsRead paper