GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction

πŸ“… 2026-08-10
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Long-term traffic scene prediction often suffers from temporal ghosting, geometric drift, and motion inconsistency, leading to unstable static structures and temporally incoherent outputs. To address this, this work proposes a training-free inference framework that enhances the outputs of pretrained video diffusion models through geometry-aware refinement during inference. The approach uniquely integrates multi-frame depth-layered rendering with vision-language model–guided view-conditioned routing, leveraging a frozen visual language model to achieve geometric stability and cross-view generalization without fine-tuning. Evaluated on the AI City Challenge Track 5, the method achieves state-of-the-art performance, significantly improving fidelity of static structures and consistency of low-level details.
πŸ“ Abstract
Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains challenging due to temporal ghosting, geometry drift, and inconsistent object motion. Recent latent video diffusion models have achieved impressive visual quality, but directly applying them to structured traffic scenes often leads to unstable geometry and degraded temporal coherence over extended horizons. We present a training-free inference framework that stabilizes reliable static structure in pretrained video predictions through multi-frame temporal context and view-conditioned routing. For front-camera videos, our method refines generated futures with a multi-frame depth-layered renderer that projects static geometry from observed history frames while preserving dynamic regions from the generative base model. For heterogeneous traffic views, a frozen vision-language model infers a coarse camera group from the observed clip and selects a specialized motion-based predictor. The framework requires neither retraining nor fine-tuning of the underlying video model and can be applied directly to pretrained generators. We validate the proposed framework on the AI City Challenge Track 5 benchmark, where our final system achieves competitive performance among the top-ranked teams. These results demonstrate that geometry-aware inference-time refinement and view-conditioned hybrid inference can improve static-geometry stability and low-level structural fidelity without changing the original model architecture.
Problem

Research questions and friction points this paper is trying to address.

future-frame prediction
geometry drift
temporal coherence
traffic scene
video generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

geometry-aware inference
training-free refinement
multi-frame depth rendering
view-conditioned routing
hybrid prediction
πŸ”Ž Similar Papers
No similar papers found.
K
Khang Minh Le
PAMI Lab, Vietnamese-German University, Vietnam
H
Hieu Dinh Trung Pham
PAMI Lab, Vietnamese-German University, Vietnam
L
Luu Thanh Danh
University of Science, Ho Chi Minh City, Vietnam
N
Nam-Tien Le
Ho Chi Minh City University of Technology, Vietnam
H
Hieu Anh Ngo
PAMI Lab, Vietnamese-German University, Vietnam
P
Phuong Huu Vu Tran
PAMI Lab, Vietnamese-German University, Vietnam
S
Son Nguyen Minh Le
PAMI Lab, Vietnamese-German University, Vietnam
N
Nguyen Trong Nghia
University of Information Technology, Ho Chi Minh City, Vietnam
T
Tu Tran Thi Cam
University of Information Technology, Ho Chi Minh City, Vietnam
H
Huy Minh Nhat Nguyen
PAMI Lab, Vietnamese-German University, Vietnam
Cuong Tuan Nguyen
Cuong Tuan Nguyen
Vietnamese-German University
neural networksmachine learningpattern recognition