FMReward: Aligning and Evaluating Audio-Driven 3D Facial Animation with Human Preferences

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between training metrics and human perception in audio-driven 3D facial animation. We introduce FMPair, the first human preference dataset for this domain, alongside FMReward, a perception-aligned reward model, and FMFL, a direct feedback fine-tuning algorithm. By leveraging pairwise preference annotation and reward modeling, our approach enables perception-oriented optimization of diffusion models. Experimental results demonstrate that FMReward significantly outperforms traditional objective metrics in predicting human preferences. Furthermore, FMFL effectively enhances the naturalness and interactive immersion of generated animations. Collectively, this work bridges the critical gap in subjective evaluation alignment research for audio-driven 3D facial animation, establishing a new paradigm for optimizing generative models based on human perceptual feedback rather than conventional quantitative measures.
📝 Abstract
Audio-driven 3D facial animation is essential for advancing immersion and interactivity in virtual experiences. Although recent advances have shown promising capabilities, the training and evaluation of existing methods typically rely on ground-truth-based errors, which fall short of aligning with human preferences. To address this, we present a comprehensive framework that learns an automatic perceptual model from human preference data and leverages it to improve and evaluate the perceptual quality of audio-driven 3D facial animation. To begin with, we construct FMPair (Facial Motion Pairwise preference), the first human preference dataset for audio-driven 3D facial animation, which is built through a systematic annotation pipeline and comprises 65,574 annotated 3D facial motion pairs from 8,834 distinct in-the-wild audio clips. Based on the pairwise comparison dataset, we propose a Facial Motion Reward model, termed FMReward, which takes audio and 3D facial motion as inputs and predicts a perceptual quality score aligned with human preferences. Building upon FMReward, we further introduce Facial Motion reward Feedback Learning (FMFL), a direct fine-tuning algorithm that leverages a pretrained reward model to optimize diffusion-based audio-driven 3D facial animation models for better alignment with human preferences. Extensive experiments demonstrate the superiority of FMReward over other metrics in aligning with human preferences and the effectiveness of FMFL in improving the perceptual quality of audio-driven 3D facial animation.
Problem

Research questions and friction points this paper is trying to address.

Audio-driven 3D facial animation
Human preference alignment
Perceptual quality evaluation
Ground-truth metrics limitation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human Preference Alignment
Facial Motion Reward Model
Reward Feedback Learning
Audio-Driven 3D Facial Animation
Perceptual Quality Evaluation
🔎 Similar Papers