In-Loop Model Adaptation with Coupled Latent-Noise Guidance for High-Fidelity Subject-Driven Text-to-Image Generation

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge in subject-driven text-to-image generation, where existing methods struggle to simultaneously preserve subject identity consistency and adapt to new visual contexts when switching reference images. The authors propose In-loop Model Adaptation (IMA), a novel approach that dynamically adjusts the diffusion model during inference without requiring pre-training or fine-tuning. By coupling latent variables with a noise-guidance mechanism and integrating DDIM inversion, masked latent consistency loss, and noise regularization, IMA enables real-time adaptation at generation time. This allows for precise retention of subject identity while maintaining strong alignment with input text prompts. Experimental results demonstrate that IMA significantly outperforms existing techniques—many of which rely on extensive training or prolonged fine-tuning—achieving state-of-the-art performance in both fidelity and subject consistency.
📝 Abstract
Text-to-image diffusion models have achieved remarkable success in generating high-quality images from a given text prompt. Subject-driven generation aims to synthesize customized images to mimic the appearance of subjects in given reference images within different visual contexts specified by the text prompts. The central challenge here is that, when the reference image changes, the diffusion model cannot efficiently adapt to different visual contexts while consistently maintaining the subject identity. Existing methods either train the model with a large domain-specific dataset or fine-tune the model using the reference image for hundreds of iterations before actual image generation. In this work, we explore a new approach, called \textit{In-Loop Model Adaptation} (IMA), which adapts the core diffusion model at each generation step during the actual process of image generation, without being trained on the reference image before the generation process. To this end, we establish a DDIM inversion chain that maps the reference image to a sequence of latent, as well as a text-to-image generation chain which generates the image from the text prompt only. We then introduce a masked latent consistency loss and a noise regularization loss to characterize the latent-noise difference between the diffusion model and these two chains at each generation step. This coupled latent-noise loss is used to guide the in-loop model adaptation to preserve the subject identity specified by the reference image while maintaining accurate alignment with the text prompt, resulting in high-fidelity text-to-image generation. Our extensive experiments demonstrate that our proposed IMA method significantly improves the performance of subject-driven text-to-image generation.
Problem

Research questions and friction points this paper is trying to address.

subject-driven generation
diffusion models
subject identity preservation
text-to-image generation
model adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-Loop Model Adaptation
Latent-Noise Guidance
Subject-Driven Generation
Diffusion Models
DDIM Inversion
🔎 Similar Papers
No similar papers found.
Y
Yushun Tang
Department of Electrical and Electronic Engineering, Southern University of Science and Technology, Shenzhen 518055, China
W
Weiming Chen
Department of Electrical and Electronic Engineering, Southern University of Science and Technology, Shenzhen 518055, China
Siyi Liu
Siyi Liu
Hong Kong University of Science and Technology (Guangzhou)
Recommender SystemsInformation Retrieval
Y
Yi Zhang
Department of Electrical and Electronic Engineering, Southern University of Science and Technology, Shenzhen 518055, China
Feng Wu
Feng Wu
National University of Singapore
Mechine LearningMedical Time Series
Zhihai He
Zhihai He
Southern University of Science and Technology
Deep learningcomputer visionmachine learningsmart cyber-physical systems