DeCoT: Decomposing Complex Instructions for Enhanced Text-to-Image Generation with Large Language Models

📅 2025-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current text-to-image (T2I) models struggle to accurately interpret complex, lengthy textual prompts, exhibiting significant semantic gaps in fine-grained detail reconstruction, spatial relationship modeling, and multi-constraint coordination. To address this, we propose DeCoT—a two-stage framework. In Stage I, a large language model (LLM) performs structured decomposition and semantic clarification of the original prompt. In Stage II, hierarchical prompt construction and chain-of-thought–driven semantic enhancement jointly model text embeddings, compositional logic, and constraint conditions. DeCoT is architecture-agnostic and seamlessly integrates with mainstream T2I models without architectural modification. Evaluated on the LongBench-T2I benchmark, DeCoT integrated with Infinity-8B achieves a score of 3.52—substantially outperforming baseline methods. Both multimodal LLM (MLLM)-based and human evaluations confirm consistent improvements across textual fidelity, spatial consistency, and compositional plausibility.

Technology Category

Application Category

📝 Abstract
Despite remarkable advancements, current Text-to-Image (T2I) models struggle with complex, long-form textual instructions, frequently failing to accurately render intricate details, spatial relationships, or specific constraints. This limitation is highlighted by benchmarks such as LongBench-T2I, which reveal deficiencies in handling composition, specific text, and fine textures. To address this, we propose DeCoT (Decomposition-CoT), a novel framework that leverages Large Language Models (LLMs) to significantly enhance T2I models' understanding and execution of complex instructions. DeCoT operates in two core stages: first, Complex Instruction Decomposition and Semantic Enhancement, where an LLM breaks down raw instructions into structured, actionable semantic units and clarifies ambiguities; second, Multi-Stage Prompt Integration and Adaptive Generation, which transforms these units into a hierarchical or optimized single prompt tailored for existing T2I models. Extensive experiments on the LongBench-T2I dataset demonstrate that DeCoT consistently and substantially improves the performance of leading T2I models across all evaluated dimensions, particularly in challenging aspects like "Text" and "Composition". Quantitative results, validated by multiple MLLM evaluators (Gemini-2.0-Flash and InternVL3-78B), show that DeCoT, when integrated with Infinity-8B, achieves an average score of 3.52, outperforming the baseline Infinity-8B (3.44). Ablation studies confirm the critical contribution of each DeCoT component and the importance of sophisticated LLM prompting. Furthermore, human evaluations corroborate these findings, indicating superior perceptual quality and instruction fidelity. DeCoT effectively bridges the gap between high-level user intent and T2I model requirements, leading to more faithful and accurate image generation.
Problem

Research questions and friction points this paper is trying to address.

Enhancing T2I models for complex long-form instructions
Improving detail accuracy in text-to-image generation
Bridging user intent with T2I model capabilities
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM decomposes complex instructions into semantic units
Multi-stage prompt integration for T2I models
Hierarchical or optimized single prompt generation
🔎 Similar Papers
No similar papers found.
X
Xiaochuan Lin
Henan Polytechnic University
X
Xiangyong Chen
Henan Polytechnic University
X
Xuan Li
Henan Polytechnic University
Y
Yichen Su
Henan Polytechnic University