Institution profile

Hong Kong Institute of Science and Technology

Academic institutionasia · hk
Official website
Research library18linked papers
Opportunities0open roles
Selected work

Representative Papers

Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation

Jul 30, 2026

This work addresses the critical privacy risks associated with releasing face training data, where existing methods struggle to disentangle identity information while preserving the class structure essential for recognition. To overcome this challenge, the authors propose a novel identity-disentangled and geometry-preserving face distillation framework that explicitly separates source identity semantics from proxy identity geometry. By enforcing orthogonal geometric preservation and aligning relational topologies, the method effectively eliminates linkability to original identities while retaining the hyperspherical proxy structure necessary for face recognition. Experimental results demonstrate that the proposed approach achieves a 3.94% improvement in TAR@FAR=1e-3 on the IJB-C surveillance benchmark, significantly outperforming baseline methods and offering a strong balance between privacy protection and model utility.

0 citationsRead paper

Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Training

Jul 02, 2026

This study investigates whether online policy self-distillation alone is sufficient to mitigate catastrophic forgetting and enable effective knowledge updating in continual post-training. Through the Self-Distillation Policy Optimization (SDPO) framework—augmented with analyses of parameter and representation space drift, as well as detection of high-frequency formatting artifacts—the authors find that while SDPO can accelerate in-domain specialization when teacher signals remain stable, dense self-distillation often induces significant representational drift. This drift, compounded by teacher-student feedback loops, amplifies formatting artifacts and degrades out-of-distribution generalization, sometimes leading to model collapse. In contrast, online reinforcement learning approaches such as GRPO adopt a more conservative update strategy, better preserving pre-existing capabilities. The work thus challenges the prevailing assumption that self-distillation inherently serves as a stable mechanism for continual learning, revealing its latent risks and delineating its operational boundaries.

0 citationsRead paper

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation

Jun 29, 2026

This work addresses the challenge of achieving flexibility, natural motion, and real-time performance in conversational talking-face generation involving multi-turn, multi-participant interactions. To this end, we propose InterTalk, a motion-driven framework that explicitly models conversational dynamics among multiple speakers. InterTalk introduces decoupled facial motion control—separating lip movements, blinking, and other expressions—and employs an iterative generation strategy enhanced by multi-source motion feedback and 3D data augmentation. For the first time, our approach enables unified, real-time (30 FPS) synthesis of highly natural talking faces with arbitrary numbers of participants across multiple dialogue turns. Extensive experiments on a newly collected large-scale multi-speaker conversational dataset demonstrate that InterTalk significantly outperforms existing methods in terms of interaction realism, generation efficiency, and overall flexibility.

0 citationsRead paper

scHelix: Asymmetric Dual-Stream Integration via Explicit Gene-Level Disentanglement

May 18, 2026

This work addresses a critical challenge in single-cell RNA sequencing data integration: the tendency of conventional whole-transcriptome harmonization approaches to over-correct, thereby compromising the preservation of genuine biological signals while removing batch effects. To overcome this limitation, the authors propose an adaptive integration framework that explicitly decouples genes at the input layer into domain-invariant Anchors and domain-sensitive Variants. The method employs an asymmetric dual-stream sparse diffusion encoder, enhanced with a stop-gradient graph cache, multi-scale structural representation learning, and a bounded residual gating mechanism to effectively prevent shortcut learning. Through an asymmetric Align-Refine-Fuse protocol, the approach consistently outperforms state-of-the-art methods across multiple benchmarks, achieving robust batch correction while faithfully retaining subtle biological clustering structures.

0 citationsRead paper

Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era

May 17, 2026

This work addresses the fundamental trade-off in continual learning for multimodal large language models between adapting to new tasks and preserving previously acquired knowledge, a challenge exacerbated when integrating reinforcement learning with verifiable rewards (RLVR) due to the absence of effective guidance mechanisms. The study introduces, for the first time, a formal notion of “reasoning transferability,” revealing the stability of reasoning-layer signals on out-of-distribution samples. Building on this insight, the authors propose Reasoning Transferability–based Dynamic Balanced Continual Learning (RDB-CL), which dynamically adjusts the strength of KL regularization at the sample level. This approach preserves reusable reasoning pathways while encouraging exploration of novel ones, thereby overcoming the limitations of conventional answer-level constraints. Experiments demonstrate that RDB-CL improves last-task accuracy by 12.0% over the original RLVR baseline, significantly outperforming existing continual learning methods.

0 citationsRead paper
Recent publications

Latest Papers

Private Face Recognition Training Dataset Publication via Identity-Decoupled and Geometry-Preserving Face Distillation

Jul 30, 2026

This work addresses the critical privacy risks associated with releasing face training data, where existing methods struggle to disentangle identity information while preserving the class structure essential for recognition. To overcome this challenge, the authors propose a novel identity-disentangled and geometry-preserving face distillation framework that explicitly separates source identity semantics from proxy identity geometry. By enforcing orthogonal geometric preservation and aligning relational topologies, the method effectively eliminates linkability to original identities while retaining the hyperspherical proxy structure necessary for face recognition. Experimental results demonstrate that the proposed approach achieves a 3.94% improvement in TAR@FAR=1e-3 on the IJB-C surveillance benchmark, significantly outperforming baseline methods and offering a strong balance between privacy protection and model utility.

0 citationsRead paper

Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Training

Jul 02, 2026

This study investigates whether online policy self-distillation alone is sufficient to mitigate catastrophic forgetting and enable effective knowledge updating in continual post-training. Through the Self-Distillation Policy Optimization (SDPO) framework—augmented with analyses of parameter and representation space drift, as well as detection of high-frequency formatting artifacts—the authors find that while SDPO can accelerate in-domain specialization when teacher signals remain stable, dense self-distillation often induces significant representational drift. This drift, compounded by teacher-student feedback loops, amplifies formatting artifacts and degrades out-of-distribution generalization, sometimes leading to model collapse. In contrast, online reinforcement learning approaches such as GRPO adopt a more conservative update strategy, better preserving pre-existing capabilities. The work thus challenges the prevailing assumption that self-distillation inherently serves as a stable mechanism for continual learning, revealing its latent risks and delineating its operational boundaries.

0 citationsRead paper

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation

Jun 29, 2026

This work addresses the challenge of achieving flexibility, natural motion, and real-time performance in conversational talking-face generation involving multi-turn, multi-participant interactions. To this end, we propose InterTalk, a motion-driven framework that explicitly models conversational dynamics among multiple speakers. InterTalk introduces decoupled facial motion control—separating lip movements, blinking, and other expressions—and employs an iterative generation strategy enhanced by multi-source motion feedback and 3D data augmentation. For the first time, our approach enables unified, real-time (30 FPS) synthesis of highly natural talking faces with arbitrary numbers of participants across multiple dialogue turns. Extensive experiments on a newly collected large-scale multi-speaker conversational dataset demonstrate that InterTalk significantly outperforms existing methods in terms of interaction realism, generation efficiency, and overall flexibility.

0 citationsRead paper

scHelix: Asymmetric Dual-Stream Integration via Explicit Gene-Level Disentanglement

May 18, 2026

This work addresses a critical challenge in single-cell RNA sequencing data integration: the tendency of conventional whole-transcriptome harmonization approaches to over-correct, thereby compromising the preservation of genuine biological signals while removing batch effects. To overcome this limitation, the authors propose an adaptive integration framework that explicitly decouples genes at the input layer into domain-invariant Anchors and domain-sensitive Variants. The method employs an asymmetric dual-stream sparse diffusion encoder, enhanced with a stop-gradient graph cache, multi-scale structural representation learning, and a bounded residual gating mechanism to effectively prevent shortcut learning. Through an asymmetric Align-Refine-Fuse protocol, the approach consistently outperforms state-of-the-art methods across multiple benchmarks, achieving robust batch correction while faithfully retaining subtle biological clustering structures.

0 citationsRead paper

Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era

May 17, 2026

This work addresses the fundamental trade-off in continual learning for multimodal large language models between adapting to new tasks and preserving previously acquired knowledge, a challenge exacerbated when integrating reinforcement learning with verifiable rewards (RLVR) due to the absence of effective guidance mechanisms. The study introduces, for the first time, a formal notion of “reasoning transferability,” revealing the stability of reasoning-layer signals on out-of-distribution samples. Building on this insight, the authors propose Reasoning Transferability–based Dynamic Balanced Continual Learning (RDB-CL), which dynamically adjusts the strength of KL regularization at the sample level. This approach preserves reusable reasoning pathways while encouraging exploration of novel ones, thereby overcoming the limitations of conventional answer-level constraints. Experiments demonstrate that RDB-CL improves last-task accuracy by 12.0% over the original RLVR baseline, significantly outperforming existing continual learning methods.

0 citationsRead paper