🤖 AI Summary
Robust facial expression capture and retargeting in unconstrained monocular videos is severely hindered by strong identity-expression coupling. Method: We propose Semantic-level Expression Representation (SEREP), the first approach to decouple identity and expression at the semantic level. Our method introduces a cycle-consistent learning framework driven by unpaired 3D data, integrated with semi-supervised domain-adaptive image-to-expression prediction—eliminating reliance on dense 3D annotations. Contributions/Results: (1) MultiREX, the first dedicated benchmark for unconstrained expression retargeting; (2) a multi-identity retargeting architecture enabling high-fidelity cross-identity expression transfer. Experiments demonstrate significant improvements over state-of-the-art methods across multiple challenging in-the-wild datasets, with markedly enhanced generalization and robustness.
📝 Abstract
Monocular facial performance capture in-the-wild is challenging due to varied capture conditions, face shapes, and expressions. Most current methods rely on linear 3D Morphable Models, which represent facial expressions independently of identity at the vertex displacement level. We propose SEREP (Semantic Expression Representation), a model that disentangles expression from identity at the semantic level. It first learns an expression representation from unpaired 3D facial expressions using a cycle consistency loss. Then we train a model to predict expression from monocular images using a novel semi-supervised scheme that relies on domain adaptation. In addition, we introduce MultiREX, a benchmark addressing the lack of evaluation resources for the expression capture task. Our experiments show that SEREP outperforms state-of-the-art methods, capturing challenging expressions and transferring them to novel identities.