Instruction-Driven 3D Facial Expression Generation and Transition
This work addresses the limitation of existing methods that typically support only six basic 3D facial expressions, hindering fine-grained and semantically driven generation and transitions. To overcome this, the authors propose the I2FET framework, which enables text-driven synthesis of arbitrary 3D facial expressions and smooth transitions between them. The key innovations include an IFED module for multimodal alignment between textual instructions and facial expression features, and a vertex reconstruction loss to enhance semantic consistency in the latent space. Evaluated on the CK+ and CelebV-HQ datasets, the proposed method significantly outperforms current approaches, generating high-fidelity, semantically accurate, and naturally continuous 3D facial expression sequences.