Towards High-Fidelity, Identity-Preserving Real-Time Makeup Transfer: Decoupling Style Generation
To address identity distortion, fairness deficits, temporal inconsistency, and non-real-time performance in real-time virtual makeup—caused by the entanglement of translucent makeup with skin tone and identity features—this paper proposes a two-stage disentanglement framework. Stage I generates pseudo-labels via unsupervised k-means clustering and graphics-based rendering, and jointly optimizes an alpha-weighted reconstruction loss and a lip-color-specific loss to achieve high-fidelity separation of makeup masks from skin tone. Stage II employs a lightweight, graphics-driven rendering module to ensure temporal coherence and fine-grained detail preservation. The method achieves real-time inference (>30 FPS) across diverse poses, expressions, and skin tones. It significantly improves transparency estimation accuracy, color fidelity, and identity preservation, outperforming state-of-the-art methods in detail reconstruction, motion smoothness, and cross-skin-tone robustness.