🤖 AI Summary
Cross-domain face reenactment suffers from poor generalization due to entanglement among identity, expression, and domain-specific style. To address this, we propose a one-shot cross-domain transfer framework that requires no paired multi-style training data. Our method employs dual encoders to disentangle domain-invariant identity features from domain-specific style features, integrates them into a diffusion-based conditional generative model, and introduces an efficient, cross-domain robust data augmentation strategy. Crucially, the framework eliminates reliance on test-time optimization or fine-tuning, enabling direct one-hop stylized expression transfer without paired cross-domain supervision. Experiments demonstrate significant improvements in identity preservation and expression fidelity across diverse visual domains—including synthetic, cartoon, and artistic styles—outperforming existing state-of-the-art methods.
📝 Abstract
Cross-domain face retargeting requires disentangled control over identity, expressions, and domain-specific stylistic attributes. Existing methods, typically trained on real-world faces, either fail to generalize across domains, need test-time optimizations, or require fine-tuning with carefully curated multi-style datasets to achieve domain-invariant identity representations. In this work, we introduce extit{StyleYourSmile}, a novel one-shot cross-domain face retargeting method that eliminates the need for curated multi-style paired data. We propose an efficient data augmentation strategy alongside a dual-encoder framework, for extracting domain-invariant identity cues and capturing domain-specific stylistic variations. Leveraging these disentangled control signals, we condition a diffusion model to retarget facial expressions across domains. Extensive experiments demonstrate that extit{StyleYourSmile} achieves superior identity preservation and retargeting fidelity across a wide range of visual domains.