Stylistic Attribute Control in Latent Diffusion Models

📅 2026-05-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing text-to-image diffusion models often compromise semantic content when performing style editing, struggling to achieve fine-grained, continuous, and disentangled control. This work proposes a method that learns disentangled editing directions from synthetic data, integrating a guidance composition mechanism with a regularized training loss and optimizing the null-text embedding to enhance DDIM inversion. This approach enables parameterized, continuous adjustment of stylistic attributes while preserving semantic consistency. Evaluated on styles such as contour emphasis, local contrast, watercolor effects, and geometric patterns, the method significantly outperforms current text-driven editing techniques, achieving more precise, coherent, and controllable style transfer.
📝 Abstract
Text-to-image diffusion models have revolutionized image synthesis and editing, but precise control over stylistic attributes remains a challenge, often causing unintended content modifications. We propose an approach for fine-grained parametric control of stylistic attributes in latent diffusion models by learning disentangled editing directions from synthetic datasets. We use guidance composition to close the domain gap between stylistically finetuned and foundation models, preserving the original image semantics while applying stylistic adjustments. To ensure consistent edits, we introduce a training regularization loss and enhance DDIM inversion with optimized null-conditional embeddings for real image editing. We validate our approach by learning from stylistically filtered synthetic datasets varying a range of stylistic attributes, including outlines, local contrast, watercolorization effects, and geometric patterns. Our evaluations demonstrate that compared to current text-based editing techniques, our method offers well-integrated, more precise and continuously adjustable stylistic modifications.
Problem

Research questions and friction points this paper is trying to address.

stylistic attribute control
latent diffusion models
image synthesis
disentangled editing
text-to-image generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

disentangled editing directions
guidance composition
DDIM inversion
null-conditional embeddings
stylistic attribute control
M
Max Reimann
Hasso-Plattner-Institute, University of Potsdam, Germany
B
Benito Buchheim
Hasso-Plattner-Institute, University of Potsdam, Germany
Jürgen Döllner
Jürgen Döllner
Professor for Computer Science, Hasso-Plattner-Institute, University of Potsdam
Computer ScienceComputer GraphicsVisualizationAnalytics