SPVC: Structured and Panoptic Video Fixing for Cross-Dataset Driving Scene Rendering

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出SPVC框架,通过结构化和全景视频修复方法解决跨数据集驾驶场景渲染中的模糊、闪烁及错位问题。
📝 Abstract
Driving scene reconstruction and rendering, especially with 3D Gaussian Splatting, has become an important component of autonomous driving simulation. However, rendered views often degrade under extrapolated ego trajectories and scene edits, producing blurry structures, temporal flicker, and foreground-background misalignment. Existing refinement methods are commonly designed for a specific setting, such as image-level novel-view repair or object-editing correction. In this paper, we introduce SPVC, a structured and panoptic video fixing framework for cross-dataset driving scene rendering. The name summarizes four design principles. (1) Structured fixing denotes the use of explicit spatial conditions, including camera pose, 3D bounding boxes, and HD maps, to guide the repair process and reduce uncontrolled hallucination. (2) Panoptic fixing refers to correcting both background rendering artifacts, such as distorted roads, buildings, and lanes, and foreground vehicle artifacts introduced by scene editing, such as inconsistent object appearance. (3) Video fixing means that the model operates on driving sequences rather than isolated frames, allowing temporal cues to be used during artifact correction. (4) Cross-dataset fixing means that a single shared network is trained and applied across multiple driving datasets, reducing the need for dataset-specific or scene-specific fixers. Concretely, we construct paired degraded-clean training data by simulating under-constrained 3DGS rendering and foreground vehicle insertion artifacts, and train a two-stage controllable video diffusion model that first addresses video-level appearance and then refines scene layout with structured controls.
Problem

Research questions and friction points this paper is trying to address.

Driving Scene Rendering
3D Gaussian Splatting
Scene Edits
Temporal Flicker
Foreground-Background Misalignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured Fixing
Panoptic Fixing
Video Fixing
Cross-dataset Fixing
🔎 Similar Papers
No similar papers found.
G
Gen Li
Institute for AI Industry Research (AIR), Tsinghua University, Beijing, China; Zhejiang University, Hangzhou, China
Shu Han
Shu Han
Yeshiva University
Information Systems
Y
Yun Xi Qiao
Tsinghua University, Beijing, China
H
Hua Chen
Great Wall Motor Company Limited, Baoding, China
X
Xuyang Dai
Great Wall Motor Company Limited, Baoding, China
B
Bohan Li
Shanghai Jiao Tong University, Shanghai, China
Hao Zhao
Hao Zhao
Tsinghua University
Computer Vision
Chaojian Li
Chaojian Li
Hong Kong University of Science and Technology
Efficient AIHardware / software codesign