Semantic-Enhanced Cross-Modal Place Recognition for Robust Robot Localization

📅 2025-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address the insufficient robustness of vision–LiDAR cross-modal place recognition under illumination, weather, and viewpoint variations in GPS-denied environments, this paper proposes SCM-PR, a semantic-enhanced cross-modal localization framework. SCM-PR innovatively adopts VMamba as the visual backbone, designs a semantic-aware feature fusion module and a semantic-guided LiDAR descriptor, and introduces a cross-modal semantic attention mechanism alongside a multi-view semantic–geometric matching strategy. Furthermore, a semantic consistency loss is proposed to enforce cross-modal semantic alignment. Evaluated on KITTI and KITTI-360, SCM-PR achieves state-of-the-art performance—significantly improving matching accuracy and robustness under complex scenes, high-resolution inputs, and large viewpoint changes.

Technology Category

Application Category

📝 Abstract
Ensuring accurate localization of robots in environments without GPS capability is a challenging task. Visual Place Recognition (VPR) techniques can potentially achieve this goal, but existing RGB-based methods are sensitive to changes in illumination, weather, and other seasonal changes. Existing cross-modal localization methods leverage the geometric properties of RGB images and 3D LiDAR maps to reduce the sensitivity issues highlighted above. Currently, state-of-the-art methods struggle in complex scenes, fine-grained or high-resolution matching, and situations where changes can occur in viewpoint. In this work, we introduce a framework we call Semantic-Enhanced Cross-Modal Place Recognition (SCM-PR) that combines high-level semantics utilizing RGB images for robust localization in LiDAR maps. Our proposed method introduces: a VMamba backbone for feature extraction of RGB images; a Semantic-Aware Feature Fusion (SAFF) module for using both place descriptors and segmentation masks; LiDAR descriptors that incorporate both semantics and geometry; and a cross-modal semantic attention mechanism in NetVLAD to improve matching. Incorporating the semantic information also was instrumental in designing a Multi-View Semantic-Geometric Matching and a Semantic Consistency Loss, both in a contrastive learning framework. Our experimental work on the KITTI and KITTI-360 datasets show that SCM-PR achieves state-of-the-art performance compared to other cross-modal place recognition methods.
Problem

Research questions and friction points this paper is trying to address.

Robust robot localization without GPS in challenging environments
Overcoming sensitivity of RGB-based methods to illumination and seasonal changes
Improving cross-modal matching between RGB images and LiDAR maps
Innovation

Methods, ideas, or system contributions that make the work stand out.

VMamba backbone extracts RGB image features
SAFF module fuses descriptors and segmentation masks
Cross-modal semantic attention improves matching accuracy
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yujia Lin
Dali University
N
Nicholas Evans
Bandırma Onyedi Eylül University