MV2: Multi-View Multi-Vehicle Driving Dataset for Novel View Synthesis
Existing novel view synthesis methods are limited in real-world driving scenarios by sparse viewpoints, dynamic objects, and single-trajectory data. To address these challenges, this work introduces a multi-view, multi-vehicle urban driving dataset—captured synchronously from cars, scooters, and drones—that enables, for the first time, large-baseline image acquisition across vehicles and trajectories, accompanied by high-precision poses and pixel-level annotations. Sequences are registered via Structure-from-Motion (SfM) and refined with manually verified correspondences to support evaluation of differentiable rendering and novel view synthesis algorithms. Comprising 12,000 images across 50 scenes, the dataset’s benchmark experiments reveal the critical impact of viewpoint span on synthesis quality and highlight performance gaps in current pose estimation methods.