Xemo-Talker: Unlock Emotions Explicitly for Audio-Driven Talking Portrait Synthesis

πŸ“… 2026-08-10
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of imprecise emotion control and the trade-off between lip synchronization and emotional expression in audio-driven talking head synthesis. We propose Xemo-Talker, a novel framework that decouples processing pathways by separating neutral speech mapping from a lightweight emotion branch. The method innovatively employs secondary principal component subspaces for emotion supervision and introduces a Tri-Loss strategy to optimize feature distribution. Experimental results demonstrate that Xemo-Talker achieves explicit and precise emotion control while maintaining high-fidelity lip synchronization and efficient inference. Notably, the model attains state-of-the-art emotion classification accuracy and generates visual quality comparable to real videos, effectively reconciling the conflict between accurate lip-syncing and expressive emotional rendering in talking head generation.
πŸ“ Abstract
Precise emotion control in audio-driven talking heads remains a challenge due to the reliance on implicit emotion regulation in existing systems, which often leads to indirect and insufficient control. Additionally, training with explicit emotion-related losses across the entire motion space poses significant difficulties due to the inherent trade-off between accurate lip synchronization and fine-grained emotion control. In this paper, we reveal a key finding: although emotional cues are distributed throughout the motion space, concentrating discriminative supervision on less-principal components achieves a better emotion-lip synchronization balance, as principal components mainly encode high-energy articulation and pose variations. Building on this insight, we propose Xemo-Talker, which first learns a neutral speech-to-motion mapping for stable articulation and lip synchronization, and then introduces a lightweight emotion branch guided by less-principal subspace supervision. To enhance emotion control, we design a Tri-Loss consisting of inter-class separation, intra-class compactness, and less-principal contrastive learning. Given an audio input, a reference image, and an emotion label, Xemo-Talker achieves state-of-the-art emotion classification accuracy while maintaining competitive lip synchronization and high inference efficiency, with performance approaching that measured on real videos.The source code is publicly available at https://github.com/chaolongy/Xemo-Talker.
Problem

Research questions and friction points this paper is trying to address.

Audio-driven talking portrait synthesis
Explicit emotion control
Lip synchronization
Emotion-lip sync trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Less-principal subspace supervision
Decoupled emotion branch
Tri-Loss
Explicit emotion control
Audio-driven talking portrait