Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the geometric instability and rendering artifacts commonly encountered in single-frame surround-view driving scene reconstruction, which stem from sparse overlap among camera views. To mitigate these issues, the authors propose the VGGD framework, which introduces visual geometric foundation model priors into this task for the first time. Specifically, a Visual Geometric Foundation Tokenizer (VGGT) generates multi-view geometric prior tokens to enhance front-end geometry modeling. A dual-path neck network is designed to disentangle geometric and appearance representations, complemented by a scale-warmup strategy and a hybrid pixel-voxel Gaussian decoder. Evaluated on the single-frame nuScenes benchmark, the proposed method significantly improves geometric consistency and rendering quality, outperforming existing approaches.
📝 Abstract
Single-frame surround-view reconstruction faces severe geometric instability and rendering artifacts due to minimal inter-camera overlap. While existing methods rely on complex decoders or auxiliary cues, they remain bottlenecked by the weak geometric capacity of upstream features. We argue that leveraging pretrained visual geometry priors strengthens upstream representations and alleviates the geometric ambiguity in sparse surround views. To this end, we propose VGGD, a visual geometry foundation-aware 3D Gaussian Splatting framework for feed-forward surround-view driving reconstruction, which shifts geometric modeling to the frontend and adapts foundation priors to the driving camera setting. First, VGGD leverages VGGT to provide transferable multi-view geometric prior tokens. Next, we introduce a Dual-Path Neck to decouple geometry-consistent and appearance-aware representations, improving appearance completion in weakly observed regions. We further apply Scale Warmup to stabilize early geometry learning and suppress scale drift under ego-pose changes. Finally, we use a hybrid pixel--volume Gaussian decoder to produce a renderable 3D Gaussian scene for novel-view synthesis. Experiments on the nuScenes single-frame benchmark show that VGGD achieves the best overall rendering quality among the compared methods and improves relative geometric consistency.
Problem

Research questions and friction points this paper is trying to address.

surround-view reconstruction
geometric instability
rendering artifacts
single-frame
3D scene reconstruction
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual geometry prior
3D Gaussian Splatting
surround-view reconstruction
geometric consistency
foundation model adaptation
🔎 Similar Papers
2024-08-29arXiv.orgCitations: 30
J
Junhong Lin
Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology, School of Electronic and Computer Engineering, Peking University, Shenzhen 518055, China
J
Jinlong Wang
Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology, School of Electronic and Computer Engineering, Peking University, Shenzhen 518055, China
Xianda Guo
Xianda Guo
PhD Student at Wuhan University
Stereo Matching, Depth Estimation,Gait Recognition
Y
Yanlun Peng
Great Wall Motor, China
W
Wei Zheng
Minieye Corporation, Shenzhen 518055, China
Guoqing Liu
Guoqing Liu
Microsoft Research AI for Science
Artificial IntelligenceReinforcement LearningLarge Language ModelsAI for Science
Hanli Wang
Hanli Wang
Tongji University
Multimedia ComputingComputer VisionImage ProcessingMachine Learning
Tiesong Zhao
Tiesong Zhao
Dept. Communication Engineering, Fuzhou University
Multimedia CommunicationVideo CodingImage Quality AssessmentHaptics
Wei Gao
Wei Gao
School of Electronic and Computer Engineering, Peking University
3D Visual Data Coding3D Visual Data ProcessingMixed RealityAI Driving & Robotics