Semantic-Enhanced Cross-Modal Place Recognition for Robust Robot Localization
To address the insufficient robustness of vision–LiDAR cross-modal place recognition under illumination, weather, and viewpoint variations in GPS-denied environments, this paper proposes SCM-PR, a semantic-enhanced cross-modal localization framework. SCM-PR innovatively adopts VMamba as the visual backbone, designs a semantic-aware feature fusion module and a semantic-guided LiDAR descriptor, and introduces a cross-modal semantic attention mechanism alongside a multi-view semantic–geometric matching strategy. Furthermore, a semantic consistency loss is proposed to enforce cross-modal semantic alignment. Evaluated on KITTI and KITTI-360, SCM-PR achieves state-of-the-art performance—significantly improving matching accuracy and robustness under complex scenes, high-resolution inputs, and large viewpoint changes.