🤖 AI Summary
This work addresses fine-grained controllable fashion image generation by proposing a garment-centric diffusion outpainting method that jointly synthesizes high-fidelity outfit display images from an input garment image, text prompts, and a face image—without explicit cloth deformation modeling. The method builds upon a latent diffusion model (LDM) to construct a multi-condition controllable generation framework. Its core contributions are: (1) a novel garment-adaptive pose prediction module enabling geometric alignment between garment and human body; (2) a multi-scale appearance customization module (MS-ACM) supporting fine-grained textual control over global style and local texture; and (3) a lightweight cross-modal feature fusion mechanism that eliminates the need for auxiliary encoders. Experiments demonstrate that our approach significantly outperforms state-of-the-art methods in garment detail fidelity, text–vision alignment accuracy, and controllability, exhibiting strong potential for commercial deployment.
📝 Abstract
In this paper, we propose a novel garment-centric outpainting (GCO) framework based on the latent diffusion model (LDM) for fine-grained controllable apparel showcase image generation. The proposed framework aims at customizing a fashion model wearing a given garment via text prompts and facial images. Different from existing methods, our framework takes a garment image segmented from a dressed mannequin or a person as the input, eliminating the need for learning cloth deformation and ensuring faithful preservation of garment details. The proposed framework consists of two stages. In the first stage, we introduce a garment-adaptive pose prediction model that generates diverse poses given the garment. Then, in the next stage, we generate apparel showcase images, conditioned on the garment and the predicted poses, along with specified text prompts and facial images. Notably, a multi-scale appearance customization module (MS-ACM) is designed to allow both overall and fine-grained text-based control over the generated model's appearance. Moreover, we leverage a lightweight feature fusion operation without introducing any extra encoders or modules to integrate multiple conditions, which is more efficient. Extensive experiments validate the superior performance of our framework compared to state-of-the-art methods.