Institution profile

University of Science and Technology Beijing

Academic institutionasia · cn
Official website
Research library331linked papers
Opportunities0open roles
Selected work

Representative Papers

LPIPS-AttnWav2Lip: Generic audio-driven lip synchronization for talking head generation in the wild

Dec 01, 2023Speech Communication

This work proposes a general-purpose approach to speaker-agnostic lip-sync generation by effectively integrating audiovisual features through a residual CBAM attention module embedded within a U-Net architecture. To enhance cross-modal alignment, a semantic alignment module is introduced to expand the receptive field, enabling precise matching between audio and visual representations. The method further incorporates the LPIPS perceptual loss to significantly improve the visual realism of generated faces and the consistency between audio and video. Experimental results demonstrate that the proposed method achieves state-of-the-art performance in both subjective evaluations and objective metrics, exhibiting strong generalization capabilities across unseen speakers and delivering high lip-sync accuracy and image quality.

7 citationsRead paper

Visual analysis of LLM-based entity resolution from scientific papers

Mar 01, 2025Visual Informatics

This paper focuses on the visual analytics support for extracting domain-specific entity from extensive scientific literature, a task with inherent limitations using traditional named entity resolution methods. With the advent of large language models (LLMs) such as GPT-4, significant improvements over conventional machine learning approaches have been achieved due to LLM's capability on entity resolution integrate abilities such as understanding multiple types of text. This research introduces a new visual analysis pipeline that integrates these advanced LLMs with versatile visualization and interaction designs to support batch entity resolution. Specifically, we focus on a specific material science field of Metal-Organic Frameworks (MOFs) and a large data collection namely CSD-MOFs. Through collaboration with domain experts in material science, we obtain well-labeled synthesis paragraphs. We propose human-in-the-loop refinement over the entity resolution process using visual analytics techniques, which allows domain experts to interactively integrate insights into LLM intelligence, including error analysis and interpretation of the retrieval-augmented generation (RAG) algorithm. Our evaluation through the case study of example selection for RAG demonstrates that this human-machine collaborative approach improved single-document entity resolution accuracy by approximately 30%.

2 citationsRead paper

Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring

Jan 20, 2026

This work addresses the high latency of chain-of-thought reasoning in multimodal large language models caused by autoregressive generation, as well as the visual information loss and hallucination induced by existing text-centric token compression methods. To this end, the authors propose V-Skip, a novel approach that introduces visual anchoring into token pruning criteria for the first time. They formulate a Visual Anchoring Information Bottleneck (VA-IB) framework augmented with a dual-path gating mechanism, which jointly evaluates linguistic surprisal and cross-modal attention flow during compression to preserve critical visual semantics. Evaluated on Qwen2-VL and Llama-3.2 series models, V-Skip achieves a 2.9× inference speedup and improves performance on DocVQA by over 30%, with negligible accuracy degradation.

1 citationsRead paper

CCL: Collaborative Curriculum Learning for Sparse-Reward Multi-Agent Reinforcement Learning via Co-evolutionary Task Evolution

May 08, 2025

In multi-agent reinforcement learning (MAS) under sparse rewards, training inefficiency and policy fragility arise from delayed feedback and difficulty in sharing experience across agents. To address these challenges, this paper proposes a collaborative curriculum learning framework. Its key contributions are: (1) a multidimensional curriculum design jointly modulating task difficulty, agent count, and environmental complexity; (2) a variational evolutionary algorithm for automated subtask generation; and (3) a co-evolutionary mechanism integrating agent policy optimization with environmental model learning. The framework unifies curriculum learning, variational evolution, MAS, and environment modeling. Evaluated on five cooperative benchmarks—including MPE and Hide-and-Seek—our method achieves significant improvements over state-of-the-art approaches: 2.1× faster convergence on average and an 18.7% increase in success rate, demonstrating both effectiveness and generalizability.

1 citationsRead paper

Global Optimality of Single-Timescale Actor-Critic under Continuous State-Action Space: A Study on Linear Quadratic Regulator

Aug 01, 2024International Joint Conference on Artificial Intelligence

This work addresses the long-standing open problem of global convergence for single-sample, single-timescale Actor-Critic algorithms in continuous state-action spaces. Using the linear quadratic regulator (LQR) as a canonical model, we establish the first global convergence guarantee to an ε-optimal policy. Our analysis integrates tools from control theory (exploiting the analytic structure of LQR), stochastic approximation, policy gradient estimation, and nonconvex optimization. We rigorously prove that the algorithm converges to an ε-optimal policy with sample complexity O(ε⁻²), matching the information-theoretic lower bound in order. This result breaks prior theoretical dependencies on discrete state-action spaces or two-timescale stepsize regimes. It provides the first tight convergence guarantee for widely deployed single-timescale Actor-Critic methods in continuous domains, thereby bridging a critical gap between theoretical analysis and practical reinforcement learning applications.

1 citationsRead paper
Recent publications

Latest Papers