LogiShot: Logically Coherent Cross-Shot Video Generation

📅 2026-08-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of narrative incoherence and visual inconsistency in cross-shot video generation caused by ambiguous user instructions. To this end, the authors propose a conditional video generation method that integrates multimodal context joint encoding with an explicit visual memory mechanism. By densely fusing video context with other conditioning signals and maintaining visual memory throughout the generation process, the approach ensures dual coherence—both semantic-logical and visual—across multiple shots. As the first study to incorporate a visual memory mechanism into conditional video generation, this method significantly outperforms existing baselines on a newly curated dataset of 110,000 samples, demonstrating particularly strong improvements on metrics evaluating cross-shot logical coherence.
📝 Abstract
Generating cross-shot videos that are logically connected is essential for content creation. Currently, most cross-shot video-generation workflows, such as short-drama production, still rely on isolated textual scripts or explicit reference images to specify the generated content. Consequently, when user instructions are underspecified or ambiguous, a generated clip may appear visually plausible on its own but fail to align with the overall narrative, leading to disjointed content. We argue that achieving cross-shot logical coherence in video generation requires establishing logical connections across shots and maintaining visual consistency. To this end, we propose LogiShot, which incorporates information through two complementary paths: 1) LogiShot jointly encodes the context video and other conditioning signals, yielding dense multimodal cues that provide visual-semantic evidence for cross-shot generation; 2) the model maintains a visual memory of the context video throughout generation to preserve visual consistency across shots. Additionally, we construct a dataset with 110K samples and a dedicated benchmark for evaluating cross-shot logical coherence. Experiments demonstrate that LogiShot consistently outperforms existing baselines in terms of logical coherence across multiple shots. Model and data will be made publicly available.
Problem

Research questions and friction points this paper is trying to address.

cross-shot video generation
logical coherence
visual consistency
narrative alignment
video generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-shot video generation
logical coherence
visual consistency
multimodal cues
visual memory
🔎 Similar Papers
No similar papers found.
S
Shuai Guo
University of Science and Technology of China
Y
Yuhang Yang
University of Science and Technology of China
Zeyu Zhang
Zeyu Zhang
Gaoling School of Artificial Intelligence, Renmin University of China
LLM-based AgentResponsible RecSysCausal Learning
Pengfei Yu
Pengfei Yu
University of Illinois at Urbana-Champaign
Natural Language ProcessingMachine Learning
W
Wei Zhai
University of Science and Technology of China
Yang Cao
Yang Cao
University of Science and Technology of China
computer visionimage processing
Z
Zheng-Jun Zha
University of Science and Technology of China