🤖 AI Summary
This work addresses large-scale, unconstrained 3D reconstruction and novel-view synthesis from sparse, uncalibrated, multi-source heterogeneous images (ground-level, surveillance, aerial) in emergency scenarios such as disaster response and law enforcement—characterized by mixed camera types, extreme illumination and viewpoint variations. Methodologically, we introduce the first real-world, cross-altitude, multi-view, uncalibrated benchmark dataset; propose an end-to-end framework integrating multi-scale NeRF, robust structure-from-motion (SfM), illumination normalization, and cross-modal geometric consistency optimization. Crucially, we decouple evaluation of camera calibration accuracy from rendering quality for the first time, thereby clarifying the true boundaries of practical challenges. Experiments demonstrate that our approach generates high-fidelity, navigable 3D models in complex real-world settings, significantly outperforming existing baselines and advancing prior-free large-scale reconstruction research.
📝 Abstract
Production of photorealistic, navigable 3D site models requires a large volume of carefully collected images that are often unavailable to first responders for disaster relief or law enforcement. Real-world challenges include limited numbers of images, heterogeneous unposed cameras, inconsistent lighting, and extreme viewpoint differences for images collected from varying altitudes. To promote research aimed at addressing these challenges, we have developed the first public benchmark dataset for 3D reconstruction and novel view synthesis based on multiple calibrated ground-level, security-level, and airborne cameras. We present datasets that pose real-world challenges, independently evaluate calibration of unposed cameras and quality of novel rendered views, demonstrate baseline performance using recent state-of-practice methods, and identify challenges for further research.