Map2Video: Street View Imagery Driven AI Video Generation

📅 2025-12-19

📈 Citations: 0

✨ Influential: 0

career value

192K/year

🤖 AI Summary

AI-driven video generation faces a fundamental bottleneck: misalignment between scene geometry and character motion spaces, resulting in poor temporal coherence and limited directorial control. To address this, we propose the first controllable video synthesis framework explicitly driven by real-world street-level geographic data (e.g., OpenStreetMap and Mapillary), integrating on-location scouting and rehearsal-style interaction into the generative pipeline—enabling map-based location selection, actor/camera placement, motion path sketching, and fine-grained lens parameter adjustment. Technically, our approach unifies Unity’s 3D engine, ComfyUI’s visual workflow interface, and the VACE video generation model to explicitly encode street-scene geometric priors and physically grounded motion constraints. In evaluations involving 12 professional filmmakers, our method significantly outperforms image-to-video baselines in spatial accuracy and scene reconstruction fidelity, while reducing cognitive load and enhancing creative controllability.

Technology Category

Application Category

📝 Abstract

AI video generation has lowered barriers to video creation, but current tools still struggle with inconsistency. Filmmakers often find that clips fail to match characters and backgrounds, making it difficult to build coherent sequences. A formative study with filmmakers highlighted challenges in shot composition, character motion, and camera control. We present Map2Video, a street view imagery-driven AI video generation tool grounded in real-world geographies. The system integrates Unity and ComfyUI with the VACE video generation model, as well as OpenStreetMap and Mapillary for street view imagery. Drawing on familiar filmmaking practices such as location scouting and rehearsal, Map2Video enables users to choose map locations, position actors and cameras in street view imagery, sketch movement paths, refine camera motion, and generate spatially consistent videos. We evaluated Map2Video with 12 filmmakers. Compared to an image-to-video baseline, it achieved higher spatial accuracy, required less cognitive effort, and offered stronger controllability for both scene replication and open-ended creative exploration.

Problem

Research questions and friction points this paper is trying to address.

Generates spatially consistent AI videos from street view imagery

Addresses shot composition, character motion, and camera control challenges

Enables location-based video creation with real-world geographic grounding

Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates Unity, ComfyUI, and VACE for video generation

Uses OpenStreetMap and Mapillary for real-world street view imagery

Enables location scouting, actor positioning, and path sketching in maps

🔎 Similar Papers

No similar papers found.

Apple

Cupertino, United States of America

AI Research Scientist, Computer Vision - Facebook Video Intelligence