Monocular Visual Simultaneous Localization and Mapping: (R)Evolution From Geometry to Deep Learning-Based Pipelines
This paper addresses the robustness bottlenecks of monocular visual SLAM in realistic scenarios—namely dynamic scenes, underwater imaging, and high-speed motion. To this end, it proposes the first comprehensive survey framework that jointly ensures classification consistency and quantifiable evaluation. The work systematically traces the evolution of geometric and end-to-end learning paradigms, unifies the modeling of these three canonical challenges, and establishes a reproducible, quantitative evaluation benchmark. It rigorously delineates the boundaries between the two paradigms, integrating multi-view geometry, nonlinear optimization, CNNs/RNNs, domain adaptation, and robust feature learning to enable cross-environment performance analysis. Empirical findings reveal distinct failure modes of existing methods under varying imaging conditions, thereby providing a systematic benchmark and an extensible research roadmap for designing and deploying robust monocular SLAM systems.