GeoFlow: Efficient Driving Video Generation via Geometry-Aligned Priors

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing driving video generation models suffer from low inference efficiency and geometric inconsistency due to their reliance on standard Gaussian noise initialization, which neglects spatiotemporal correlations across frames. To address this, this work proposes a Geometry-Aligned Prior (GAP) mechanism that replaces conventional Gaussian noise with a spatially adaptive initial state derived from multi-view geometric constraints, thereby explicitly injecting geometric information at the onset of generation. This approach substantially shortens the required sampling trajectory, enabling significant improvements in few-step generation quality with only a few hours of fine-tuning under either flow-matching or diffusion frameworks. When fully trained, GAP markedly reduces the number of inference steps needed to achieve state-of-the-art performance while simultaneously enhancing temporal geometric consistency and overall generation efficiency.
📝 Abstract
Generative models like Diffusion Models and Flow Matching have demonstrated remarkable capabilities in synthesizing high-fidelity driving videos, but are severely constrained by high inference latency due to the requirement of extensive sampling steps. We argue that this inefficiency stems from the prevailing reliance on a standard Gaussian source distribution, where consecutive frames are initialized as independent Gaussian noise. This paradigm disregards the rich spatiotemporal correlations inherent in driving videos, compelling the model to regenerate deterministic scene structures existing in previous frames from noise, which is both computationally redundant and prone to geometric inconsistency. To address this problem, we propose GeoFlow, a novel framework designed to achieve efficient driving video generation by harnessing explicit geometric priors. Instead of sampling from standard Gaussian noise, we leverage multi-view geometry and spatially-adaptive noise injection to construct a Geometry-Aligned Prior (GAP) distribution as starting point. This initialization bridges the gap between source distribution and data distribution, yielding a significantly straighter and shorter sampling trajectory. Extensive experiments demonstrate that GeoFlow can achieve remarkable efficiency of both training and inference: merely several hours of fine-tuning on baseline models can significantly boost few-step generation quality, while fully converged training drastically reduces number of inference steps required for state-of-the-art video generation.
Problem

Research questions and friction points this paper is trying to address.

driving video generation
inference latency
spatiotemporal correlation
geometric inconsistency
sampling efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geometry-Aligned Prior
Efficient Video Generation
Multi-view Geometry
Spatially-Adaptive Noise
Driving Video Synthesis