StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing generative previsualization approaches rely on one-shot image or video synthesis, which limits fine-grained, iterative control over scene structure, motion, camera parameters, and spatiotemporal dynamics. This work proposes a controllable and editable previsualization framework centered on an explicit, persistent 3D world state. It introduces a three-stage pipeline—construction, evolution, and access—featuring a novel prior-guided, conflict-aware dual-view initialization to build the 3D state. Structured state transitions enable user-intent-driven local edits and state reuse, eliminating the need for full-scene regeneration. Off-the-shelf video generation models are leveraged to enhance visual fidelity. The approach substantially improves both editing flexibility and visual quality, offering a powerful solution for high-quality, interactive video prototyping in film, gaming, and related creative domains.
📝 Abstract
Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras, and spatial-temporal dynamics. Yet existing generative methods rely on simple prompts to jointly control all of these factors through one-shot image or video synthesis, offering weak controllability and limited support for iterative editing. Fundamentally, a world comprises multiple elements with geometry, appearance, and other attributes, together with cameras. Different frames are produced through local modifications or recombinations of this shared state, which is otherwise largely reused. Therefore, we argue that the missing component is an explicit and persistent working state. To address this, we present StateFlow, a state-centric framework for generative previsualization. Rather than generating videos in one shot, StateFlow uses an editable 3D world to organize scene structure, evolution, and cameras, while off-the-shelf video models enhance visual quality when higher fidelity is desired. This world is maintained as a persistent structured 3D state of scene elements and camera configurations, serving as the core working representation for previsualization. Built on this insight, StateFlow has three stages to construct, evolve, and access the world state. State construction lifts generated 2D content into a coherent 3D world through prior-guided, conflict-aware dual-view initialization, while State evolution translates user intent into structured state transitions while preserving world memory, avoiding full-scene regeneration for each edit. State access uses render-feedback reflection to refine camera plans into visually feasible trajectories, avoiding reliance on VLM semantics alone. Experiments show that StateFlow produces high-quality 3D worlds for video creation and game-like prototyping.
Problem

Research questions and friction points this paper is trying to address.

previsualization
3D world state
controllability
iterative editing
generative modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

state-centric framework
persistent 3D world state
structured state transition
render-feedback reflection
generative previsualization
🔎 Similar Papers
No similar papers found.
Yuyang Yin
Yuyang Yin
Beijing Jiaotong University
Computer VisionAIGC
Zixiang Li
Zixiang Li
Beijing Jiaotong University
L
Longxuan Deng
Beijing Jiaotong University
H
Hongkai Li
Beijing Jiaotong University
S
Shifang Zhao
Beijing Jiaotong University
J
Junnan Liu
Beijing Jiaotong University
W
Weirong Huang
Beijing Jiaotong University
M
Mengyu Wang
Beijing Jiaotong University
T
Tianxiao Fu
Mootion AI
Y
Yikai Wang
Beijing Normal University
Peng-Shuai Wang
Peng-Shuai Wang
Assistant Professor, Peking University
Geometry processing3D Deep LearningComputer graphics
X
Xiaojie Jin
Beijing Jiaotong University
Y
Yao Zhao
Beijing Jiaotong University
Yunchao Wei
Yunchao Wei
Professor, Beijing Jiaotong University, UTS, UIUC, NUS
Computer VisionMachine Learning