Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation

📅 2026-08-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing methods struggle to balance controllability and visual fidelity: implicit representations often produce averaged expressions due to insufficient structural guidance, while explicit geometric approaches lose high-frequency textural details. To address this, this work proposes GemTalk, a novel framework that synergistically integrates implicit affective semantic features with explicit Blendshape geometric priors. Built upon a diffusion model, GemTalk employs a Video-Audio Emotion Perception (V-AEP) module to extract multimodal emotional cues, a Dynamic Geometry-Prior Generator (D-GPG) to produce identity-aware Blendshape coefficients, and a Geometry-guided Emotion Modulation (GEM) module that recalibrates the intensity of implicit features using geometric information for precise, continuous control over emotional expression. Experiments demonstrate that GemTalk significantly outperforms state-of-the-art methods in both expressive dynamics and photorealistic quality, achieving high-fidelity, highly controllable emotional talking-face generation.
📝 Abstract
Audio-driven emotional talking face generation aims to synthesize realistic videos with expressive facial dynamics. However, existing methods struggle to balance controllability and visual fidelity. Although implicit representations capture rich semantics, they lack structural guidance, often resulting in averaged emotional expressions. In contrast, explicit geometric methods offer better control over facial expressions but tend to sacrifice high-frequency texture details. To address it, we propose GemTalk, a diffusion-based framework that combines the semantic richness of implicit representations with the structural precision of explicit geometric priors. We introduce a Vision-guided Audio Emotion Projection (V-AEP) module to extract implicit emotional lip and expression features. At the same time, a Diffusion-based Geometric Priors Generator (D-GPG) generates identity-aware blendshape coefficients as explicit structural priors. Crucially, our Geometry-guided Emotion Modulation (GEM) module leverages these geometric priors to recalibrate the magnitude of implicit features, enabling precise, continuous control over emotional expressions, especially emotion intensity, without sacrificing visual quality. Extensive experiments show GemTalk achieves superior performance in photo-realism, and facial emotional dynamics.
Problem

Research questions and friction points this paper is trying to address.

emotional talking face generation
controllability
visual fidelity
implicit representations
explicit geometric priors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geometry-guided Emotion Modulation
Diffusion-based Talking Face
Implicit-Explicit Fusion
Emotion Intensity Control
Photorealistic Facial Animation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chenggong Hu
School of Software Technology, Zhejiang University
S
Shaoyin Ma
School of Software Technology, Zhejiang University
Yi Wang
Yi Wang
Ph.D. Student of CS, ZJU
Generative ModelMultimodal Foundation Model
Li Sun
Li Sun
Ningbo Innovation Center, Zhejiang University
computer visionroboticsmachine learningartificial intelligence
M
Mingli Song
College of Computer Science and Technology, Zhejiang University
Jie Song
Jie Song
Professor, University of Massachusetts Chan Medical School
biomaterialsregenerative medicine