🤖 AI Summary
This study addresses the limitations of discrete emotion labels and audio-visual temporal misalignment in emotional 3D talking head generation by proposing an audio-driven framework based on continuous Valence-Arousal (VA) space. Through dynamic emotion modulation, multi-scale temporal modeling, and an adaptive fusion mechanism, the method effectively decouples high-frequency articulation from low-frequency emotional dynamics. Furthermore, we introduce 3D-VA-MEAD, the first large-scale 3D emotional dataset with automatic VA annotations. Experimental results demonstrate that our approach outperforms state-of-the-art methods in both lip-sync accuracy and emotional naturalness while enabling fine-grained, smooth emotion control, thereby validating the efficacy of continuous emotion modeling for expressive 3D avatar synthesis.
📝 Abstract
Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing methods often rely on discrete emotion categories, which fail to capture the continuous evolution of affect. They also overlook the temporal frequency mismatch between audio articulation and emotional expression. In this paper, we propose CETalk, an audio-driven 3D facial animation framework conditioned on continuous Valence--Arousal (VA) representations for fine-grained emotion control. CETalk predicts a sequence of FLAME parameters through three key components: a Dynamic Emotion Modulation Module that adaptively scales emotional intensity using audio-derived cues; a Multi-Scale Temporal Modeling mechanism that employs parallel branches to decouple high-frequency articulatory movements from low-frequency emotional dynamics; and a Dynamic Fusion Mechanism that integrates these multi-scale features via an adaptive gating network. To support training and evaluation, we construct 3D-VA-MEAD, a large-scale dataset with automatically estimated VA annotations and reconstructed 3D facial motions. Extensive experiments demonstrate that CETalk outperforms state-of-the-art methods in both lip-sync accuracy and emotional expressiveness, while enabling smooth and controllable emotion transitions.