Institution profile

South China University of Technology

Academic institutionasia · cn
Official website
Research library964linked papers
Opportunities0open roles
Selected work

Representative Papers

Model Adaptation: Unsupervised Domain Adaptation Without Source Data

Jun 01, 2020Computer Vision and Pattern Recognition

This paper addresses unsupervised model adaptation from a source domain to a target domain without access to any source-domain data or labels—termed *source-free* adaptation—aiming solely to improve the generalization of a pre-trained source model on unlabeled target data. To this end, we propose a Collaborative Class-Conditional Generative Adversarial Network (CC-GAN) framework: it models the semantic structure of the target domain via class-conditional generation; enforces weight constraints derived from the source model to preserve its discriminative capability; and incorporates clustering-driven feature regularization to enhance the discriminability of target-domain representations. Evaluated across multiple cross-domain vision tasks, our method achieves significant performance gains over conventional source-dependent adaptation approaches—using only unlabeled target data. It is the first to empirically validate the effectiveness, robustness, and scalability of model adaptation under the source-free setting.

464 citations53 influentialRead paper

Never-Ending Behavior-Cloning Agent for Robotic Manipulation

Mar 01, 2024

Embodied robots struggle with 3D scene understanding and human-level task generalization in unstructured environments due to reliance on multimodal observations. Method: This paper proposes a lifelong language-conditioned behavioral cloning framework tailored for real-world scenarios. It introduces the first lifelong behavioral cloning paradigm; designs a skill-sharing semantic rendering and representation distillation module to mitigate 3D representation blind spots; and develops a skill-specific evolutionary planner enabling human-like incremental knowledge embedding in a low-rank latent space. Contribution/Results: Evaluated on a newly established lifelong manipulation benchmark, the method significantly outperforms state-of-the-art approaches. The code, dataset, and visualization results are publicly released, demonstrating strong cross-task sequential adaptability and robustness to continual learning.

5 citationsRead paper

SCALA: Split Federated Learning with Concatenated Activations and Logit Adjustments

May 08, 2024arXiv.org

To address label distribution skew arising from data heterogeneity and partial client participation in Split Federated Learning (SFL), this paper proposes a novel collaborative logit calibration framework. The method introduces (1) a cross-client activation concatenation mechanism—first of its kind—to enable server-side unified modeling of the global label distribution; and (2) a dual-sided logit adjustment strategy operating jointly at the server and client levels, explicitly correcting class-wise bias in local loss functions—an innovation not previously realized in SFL. Theoretical analysis establishes enhanced convergence robustness under non-IID data assumptions. Empirical evaluation across multiple benchmark datasets demonstrates an average accuracy improvement of 3.2% and significantly accelerated convergence compared to state-of-the-art baselines.

3 citationsRead paper

How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing

Feb 02, 2026

This work addresses the limitation of existing image editing models and evaluation benchmarks, which predominantly rely on textual instructions and struggle to support visual directives—such as sketches—that are integral to human multimodal interaction. To bridge this gap, we introduce VIBE, the first systematic benchmark for vision-instructed image editing, which defines a three-tiered hierarchy of task complexity ranging from referential localization and shape manipulation to causal reasoning. We further develop a fine-grained automatic evaluation framework based on multimodal large language models (LMMs) and use it to assess 17 open- and closed-source models across diverse visual instructions. Our evaluation reveals that closed-source models exhibit初步 stronger instruction-following capabilities, yet all models suffer significant performance degradation on higher-order tasks, highlighting critical limitations and pointing toward promising directions for future research.

2 citationsRead paper
Recent publications

Latest Papers