Audit & Repair: An Agentic Framework for Consistent Story Visualization in Text-to-Image Diffusion Models

📅 2025-06-23

📈 Citations: 0

✨ Influential: 0

career value

186K/year

🤖 AI Summary

Text-to-image diffusion models struggle to maintain cross-frame consistency of characters and objects in multi-panel story generation, undermining narrative coherence. To address this, we propose a multi-agent collaborative audit-and-repair framework that enables iterative, fine-grained panel-level correction without regenerating the entire sequence. The framework decouples auditing and editing modules, ensuring compatibility with diverse diffusion architectures—including Flux and Stable Diffusion. During inference, agents jointly detect visual inconsistencies (e.g., identity mismatches, attribute drift) and apply targeted edits to restore fidelity. Experiments demonstrate substantial improvements over state-of-the-art baselines across quantitative metrics—including CLIP-IoU and identity similarity—as well as qualitative human evaluation. Our approach significantly enhances inter-panel visual consistency and character persistence, establishing a novel paradigm for controllable long-sequence image generation.

Technology Category

Application Category

📝 Abstract

Story visualization has become a popular task where visual scenes are generated to depict a narrative across multiple panels. A central challenge in this setting is maintaining visual consistency, particularly in how characters and objects persist and evolve throughout the story. Despite recent advances in diffusion models, current approaches often fail to preserve key character attributes, leading to incoherent narratives. In this work, we propose a collaborative multi-agent framework that autonomously identifies, corrects, and refines inconsistencies across multi-panel story visualizations. The agents operate in an iterative loop, enabling fine-grained, panel-level updates without re-generating entire sequences. Our framework is model-agnostic and flexibly integrates with a variety of diffusion models, including rectified flow transformers such as Flux and latent diffusion models such as Stable Diffusion. Quantitative and qualitative experiments show that our method outperforms prior approaches in terms of multi-panel consistency.

Problem

Research questions and friction points this paper is trying to address.

Maintaining visual consistency in multi-panel story visualization

Preserving key character attributes across generated panels

Correcting inconsistencies without regenerating entire sequences

Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-agent framework for visual consistency

Iterative loop for panel-level updates

Model-agnostic integration with diffusion models

🔎 Similar Papers

DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion