CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of discrete emotion labels and audio-visual temporal misalignment in emotional 3D talking head generation by proposing an audio-driven framework based on continuous Valence-Arousal (VA) space. Through dynamic emotion modulation, multi-scale temporal modeling, and an adaptive fusion mechanism, the method effectively decouples high-frequency articulation from low-frequency emotional dynamics. Furthermore, we introduce 3D-VA-MEAD, the first large-scale 3D emotional dataset with automatic VA annotations. Experimental results demonstrate that our approach outperforms state-of-the-art methods in both lip-sync accuracy and emotional naturalness while enabling fine-grained, smooth emotion control, thereby validating the efficacy of continuous emotion modeling for expressive 3D avatar synthesis.
📝 Abstract
Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing methods often rely on discrete emotion categories, which fail to capture the continuous evolution of affect. They also overlook the temporal frequency mismatch between audio articulation and emotional expression. In this paper, we propose CETalk, an audio-driven 3D facial animation framework conditioned on continuous Valence--Arousal (VA) representations for fine-grained emotion control. CETalk predicts a sequence of FLAME parameters through three key components: a Dynamic Emotion Modulation Module that adaptively scales emotional intensity using audio-derived cues; a Multi-Scale Temporal Modeling mechanism that employs parallel branches to decouple high-frequency articulatory movements from low-frequency emotional dynamics; and a Dynamic Fusion Mechanism that integrates these multi-scale features via an adaptive gating network. To support training and evaluation, we construct 3D-VA-MEAD, a large-scale dataset with automatically estimated VA annotations and reconstructed 3D facial motions. Extensive experiments demonstrate that CETalk outperforms state-of-the-art methods in both lip-sync accuracy and emotional expressiveness, while enabling smooth and controllable emotion transitions.
Problem

Research questions and friction points this paper is trying to address.

Emotional 3D talking head generation
Continuous emotion control
Valence-Arousal
Temporal frequency mismatch
Audio-driven facial animation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continuous Valence-Arousal
Dynamic Emotion Modulation
Multi-Scale Temporal Modeling
Audio-Driven 3D Talking Head
3D-VA-MEAD Dataset
🔎 Similar Papers
2024-03-19IEEE Workshop/Winter Conference on Applications of Computer VisionCitations: 4