DRDM: A Disentangled Representations Diffusion Model for Synthesizing Realistic Person Images

📅 2024-12-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing virtual try-on and portrait editing methods suffer from limb distortion, loss of fine details, and clothing style degradation during pose transfer. To address these issues, we propose the Decoupled Representation Diffusion Model (DRDM), featuring two key innovations: (1) a Body Subspace Decoupling Block (BSDB) that explicitly disentangles pose and appearance representations at the body-part level, and (2) a parsing-map-driven classifier-free guidance sampling mechanism for precise semantic control. DRDM integrates diffusion modeling, pose-aware encoding, part-wise feature decoupling, and semantic parsing map conditioning. Evaluated on DeepFashion, DRDM achieves state-of-the-art performance in pose fidelity and appearance controllability—demonstrating superior limb structural accuracy, texture detail realism, and clothing style consistency compared to prior approaches.

Technology Category

Application Category

📝 Abstract
Person image synthesis with controllable body poses and appearances is an essential task owing to the practical needs in the context of virtual try-on, image editing and video production. However, existing methods face significant challenges with details missing, limbs distortion and the garment style deviation. To address these issues, we propose a Disentangled Representations Diffusion Model (DRDM) to generate photo-realistic images from source portraits in specific desired poses and appearances. First, a pose encoder is responsible for encoding pose features into a high-dimensional space to guide the generation of person images. Second, a body-part subspace decoupling block (BSDB) disentangles features from the different body parts of a source figure and feeds them to the various layers of the noise prediction block, thereby supplying the network with rich disentangled features for generating a realistic target image. Moreover, during inference, we develop a parsing map-based disentangled classifier-free guided sampling method, which amplifies the conditional signals of texture and pose. Extensive experimental results on the Deepfashion dataset demonstrate the effectiveness of our approach in achieving pose transfer and appearance control.
Problem

Research questions and friction points this paper is trying to address.

Human Pose Control
Body Part Deformation
Clothing Style Inconsistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decoupled Representation
Pose and Appearance Control
Body-aware Diffusion Model
🔎 Similar Papers
Nanning Normal University | Zhidayuan AI Lab | South China Agricultural University | Sun Yat-sen University
E
Enbo Huang
Guangxi Key Lab of Human-machine Interaction and Intelligent Decision, Nanning Normal University, Nanning, China
Y
Yuan Zhang
Zhidayuan AI Lab, Nanning, China
Faliang Huang
Faliang Huang
GXHIID, Nanning Normal University
Human-Machine InteractionArtificial Intelligence
Guangyu Zhang
Guangyu Zhang
College of Mathematics and Informatics, South China Agricultural University, Guangzhou, China
Y
Yang Liu
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China