Institution profile

Modiface

Industry researchnorthamerica · ca
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Towards High-Fidelity, Identity-Preserving Real-Time Makeup Transfer: Decoupling Style Generation

Sep 02, 2025

To address identity distortion, fairness deficits, temporal inconsistency, and non-real-time performance in real-time virtual makeup—caused by the entanglement of translucent makeup with skin tone and identity features—this paper proposes a two-stage disentanglement framework. Stage I generates pseudo-labels via unsupervised k-means clustering and graphics-based rendering, and jointly optimizes an alpha-weighted reconstruction loss and a lip-color-specific loss to achieve high-fidelity separation of makeup masks from skin tone. Stage II employs a lightweight, graphics-driven rendering module to ensure temporal coherence and fine-grained detail preservation. The method achieves real-time inference (>30 FPS) across diverse poses, expressions, and skin tones. It significantly improves transparency estimation accuracy, color fidelity, and identity preservation, outperforming state-of-the-art methods in detail reconstruction, motion smoothness, and cross-skin-tone robustness.

0 citationsRead paper

S$^2$Edit: Text-Guided Image Editing with Precise Semantic and Spatial Control

Jul 06, 2025

Existing text-guided image editing methods based on diffusion models suffer from coarse semantic control, inaccurate spatial localization, loss of subject identity and high-frequency details, and erroneous modifications to irrelevant regions due to semantic entanglement. To address these issues, we propose a Semantic-Spatial Dual-Decoupling Editing framework. Methodologically: (1) we introduce learnable text token embeddings to explicitly represent subject identity and impose feature-space orthogonality constraints to decouple identity from attribute semantics; (2) we incorporate object-mask-guided cross-attention to achieve spatially focused editing. Our approach requires only lightweight adaptation—no full-model fine-tuning—on pretrained text-to-image diffusion models. Extensive qualitative and quantitative evaluations demonstrate significant improvements over state-of-the-art methods in editing accuracy, identity preservation, and detail fidelity. Moreover, our framework successfully supports complex multi-attribute editing tasks, such as makeup transfer.

0 citationsRead paper
Recent publications

Latest Papers

Towards High-Fidelity, Identity-Preserving Real-Time Makeup Transfer: Decoupling Style Generation

Sep 02, 2025

To address identity distortion, fairness deficits, temporal inconsistency, and non-real-time performance in real-time virtual makeup—caused by the entanglement of translucent makeup with skin tone and identity features—this paper proposes a two-stage disentanglement framework. Stage I generates pseudo-labels via unsupervised k-means clustering and graphics-based rendering, and jointly optimizes an alpha-weighted reconstruction loss and a lip-color-specific loss to achieve high-fidelity separation of makeup masks from skin tone. Stage II employs a lightweight, graphics-driven rendering module to ensure temporal coherence and fine-grained detail preservation. The method achieves real-time inference (>30 FPS) across diverse poses, expressions, and skin tones. It significantly improves transparency estimation accuracy, color fidelity, and identity preservation, outperforming state-of-the-art methods in detail reconstruction, motion smoothness, and cross-skin-tone robustness.

0 citationsRead paper

S$^2$Edit: Text-Guided Image Editing with Precise Semantic and Spatial Control

Jul 06, 2025

Existing text-guided image editing methods based on diffusion models suffer from coarse semantic control, inaccurate spatial localization, loss of subject identity and high-frequency details, and erroneous modifications to irrelevant regions due to semantic entanglement. To address these issues, we propose a Semantic-Spatial Dual-Decoupling Editing framework. Methodologically: (1) we introduce learnable text token embeddings to explicitly represent subject identity and impose feature-space orthogonality constraints to decouple identity from attribute semantics; (2) we incorporate object-mask-guided cross-attention to achieve spatially focused editing. Our approach requires only lightweight adaptation—no full-model fine-tuning—on pretrained text-to-image diffusion models. Extensive qualitative and quantitative evaluations demonstrate significant improvements over state-of-the-art methods in editing accuracy, identity preservation, and detail fidelity. Moreover, our framework successfully supports complex multi-attribute editing tasks, such as makeup transfer.

0 citationsRead paper