Concept Guidance: Precise, Training-Free Latent Control for Text-to-Image Generation

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of text-to-image models regarding continuous concept guidance and unreliable local structure generation by proposing a training-free, gradient-free method for precise concept control. By quantifying inter-layer concept mutual information to reveal localization properties in generation, the approach employs a skip-layer prediction weighted combination strategy to achieve accurate denoising guidance within the latent space. Experimental results demonstrate that this method significantly enhances multi-object generation performance and local consistency across mainstream architectures, including PixArt-alpha, SD3, and FLUX.1. Consequently, this work establishes an efficient and reliable paradigm for concept-level regulation in text-to-image synthesis, overcoming critical challenges in fine-grained semantic control without requiring additional model retraining or gradient computation.
📝 Abstract
Text-to-image diffusion models have two major drawbacks that severely limit their practical utility: (1) standard models lack an intrinsic mechanism for continuous, concept-specific guidance (e.g., for precisely controlling how aesthetically pleasing an image looks), and (2) they lack reliability for tasks requiring high local coherence (e.g., generating text or human hands). To tackle these issues, we introduce a novel notion of concept-wise mutual information and find large, concept-dependent differences between individual layers, demonstrating that the generation of specific structures is localized in distinct parts of the network. We exploit this insight by reinforcing the impact of concept-relevant layers in Concept Guidance (CoG), a precise, target-specific guidance method that works for models out-of-the-box without additional training, external models, gradients, or prompt engineering. CoG first quantifies each layer's concept-specific impact and then guides denoising using a weighted combination of predictions generated with concept-relevant layers skipped. We demonstrate performance increases across various targets and popular models like PixArt-alpha, SD3, SD3.5, and FLUX.1-dev. Code is available at https://github.com/CompVis/concept_guidance
Problem

Research questions and friction points this paper is trying to address.

Text-to-Image Generation
Concept Guidance
Local Coherence
Diffusion Models
Latent Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Concept Guidance
Training-Free
Concept-wise Mutual Information
Latent Control
Text-to-Image Generation
🔎 Similar Papers
No similar papers found.
N
Nikolai Röhrich
LMU Munich, Germany; Konrad Zuse School of Excellence in Reliable AI (relAI), Germany
I
Isabell Hans
LMU Munich, Germany; Konrad Zuse School of Excellence in Reliable AI (relAI), Germany
Felix Krause
Felix Krause
PhD Student, LMU Munich
Deep LearningComputer Vision
Björn Ommer
Björn Ommer
Professor, Computer Vision & Learning Group (CompVis), University of Munich
computer visionmachine learningartificial intelligencecognitive science