Hybrid Rendering for Multimodal Autonomous Driving: Merging Neural and Physics-Based Simulation

๐Ÿ“… 2025-03-12
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
To address the limited generalization and poor physical controllability of neural reconstruction in autonomous driving simulation, this paper proposes a hybrid framework integrating neural and physics-based rendering. Methodologically, we introduce NeRF2GSโ€”a novel distillation paradigm where a depth-regularized Neural Radiance Field (NeRF) serves as the teacher model and 3D Gaussian Splatting (3DGS) as the student. We incorporate LiDAR point cloudโ€“robust depth supervision, block-parallel optimization, and depth-aware compositing to enable reconstruction of scenes spanning hundreds of square kilometers, while jointly outputting RGB, semantic segmentation, surface normals, and depth maps. Our contributions include the first demonstration of arbitrary placement of dynamic agents, adjustable environmental parameters, and multi-view real-time rendering (>30 FPS). This significantly improves novel-view synthesis quality for road surfaces and lane markings, and ensures compatibility with multi-camera setups and high-fidelity LiDAR simulation.

Technology Category

Application Category

๐Ÿ“ Abstract
Neural reconstruction models for autonomous driving simulation have made significant strides in recent years, with dynamic models becoming increasingly prevalent. However, these models are typically limited to handling in-domain objects closely following their original trajectories. We introduce a hybrid approach that combines the strengths of neural reconstruction with physics-based rendering. This method enables the virtual placement of traditional mesh-based dynamic agents at arbitrary locations, adjustments to environmental conditions, and rendering from novel camera viewpoints. Our approach significantly enhances novel view synthesis quality -- especially for road surfaces and lane markings -- while maintaining interactive frame rates through our novel training method, NeRF2GS. This technique leverages the superior generalization capabilities of NeRF-based methods and the real-time rendering speed of 3D Gaussian Splatting (3DGS). We achieve this by training a customized NeRF model on the original images with depth regularization derived from a noisy LiDAR point cloud, then using it as a teacher model for 3DGS training. This process ensures accurate depth, surface normals, and camera appearance modeling as supervision. With our block-based training parallelization, the method can handle large-scale reconstructions (greater than or equal to 100,000 square meters) and predict segmentation masks, surface normals, and depth maps. During simulation, it supports a rasterization-based rendering backend with depth-based composition and multiple camera models for real-time camera simulation, as well as a ray-traced backend for precise LiDAR simulation.
Problem

Research questions and friction points this paper is trying to address.

Enhances novel view synthesis for autonomous driving simulations.
Combines neural and physics-based rendering for dynamic environments.
Supports large-scale reconstructions and real-time rendering capabilities.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Combines neural and physics-based rendering techniques
Uses NeRF2GS for enhanced view synthesis quality
Supports large-scale reconstructions and real-time rendering
๐Ÿ”Ž Similar Papers
๐Ÿ’ผ Related Jobs
No related jobs found.
M
M'at'e T'oth
aiMotive
P
P'eter Kov'acs
aiMotive
Z
Zolt'an Bendefy
aiMotive
Z
Zolt'an Hortsin
aiMotive
B
Bal'azs Ter'eki
aiMotive
T
Tam'as Matuszka
aiMotive