Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation
This work addresses the limitations of existing speech-driven facial animation methods, which predominantly rely on monolingual data and struggle to accommodate linguistic variations and individual speaking styles in multilingual settings. The authors propose a unified diffusion model architecture that implicitly encodes language information through text embeddings and extracts stylistic representations from reference facial sequences, enabling personalized multilingual facial animation without requiring predefined language or speaker labels. Notably, this approach is the first to jointly model the interactive effects of language and speaking style, facilitating cross-lingual and cross-speaker generalization under label-free conditions. Experimental results demonstrate that the method outperforms current state-of-the-art approaches in both monolingual and multilingual scenarios, producing animations that exhibit more natural and realistic articulatory timing, habitual facial gestures, and temporal coherence.