Sidecar: Training-Free Semantic Reuse for Character-Consistent Free-form Visual Storytelling

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决自由形式视觉叙事中角色一致性问题,提出Sidecar模块,无需额外训练,通过语义增强保留初始描述信息,提高图像与提示及角色一致性。
📝 Abstract
Visual storytelling requires generating images that follow a narrative while preserving consistent character identities across frames. In free-form story generation, a character is fully described only when first introduced and is later referred to by a type-level mention or pronoun. Although this setting better reflects natural storytelling, later prompts may omit important identity-related semantics, making character consistency more difficult to maintain. We propose \textbf{Sidecar}, a plug-and-play semantic augmentation module that preserves entity-level information from the initial description and injects the missing semantics into later prompt embeddings. Sidecar requires no additional training and does not modify the architecture of the base diffusion model. Experiments on FreeStoryBench show that Sidecar consistently improves prompt-image alignment and character consistency across multiple SDXL- and FLUX-based baselines, with negligible computational overhead.
Problem

Research questions and friction points this paper is trying to address.

Visual Storytelling
Character Consistency
Free-form Story Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-Free
Semantic Reuse
Character-Consistent
Free-form Visual Storytelling
Plug-and-Play