DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出DXPR框架,通过将相机图像和LiDAR扫描转换为统一的深度图像表示,并利用视觉基础模型实现跨模态地点识别,解决了机器人在不同环境条件下使用相机定位的问题。
📝 Abstract
We present DXPR, a depth-based cross-modal place recognition (CMPR) framework that uses vision foundation models (VFMs) to match monocular camera queries against a LiDAR map without modality-specific encoders. This enables robots and autonomous vehicles to robustly localize using only cameras within pre-built LiDAR maps, even under severe seasonal, weather, and illumination changes. The key idea is to convert both camera images and LiDAR scans into a unified depth image representation so that a single VFM backbone with an aggregation head can learn modality-invariant global descriptors. To make pairwise metric learning faithful to scene geometry, we introduce a geometry-aware overlap miner: after cross-modal scale alignment of camera and LiDAR depth, we forward-warp measurements between views to compute a pixel-level overlap score. This score relabels ambiguous pairs and adaptively modulates the positive margin in a multi-similarity loss to avoid overfitting on weakly overlapping views. Extensive experiments on KITTI odometry and Boreas demonstrate strong performance and robustness across seasons, weather, and day/night. On KITTI, DXPR achieves near-perfect Recall@1 on most sequences and outperforms prior CMPR baselines. On Boreas, DXPR achieves intra-sequence performance on par with a strong single-modal baseline (DINOv2-SALAD), while showing clear improvements in the more challenging inter-sequence setting. Compared with RangeBEV, our method consistently performs better in both intra- and inter-sequence evaluations, demonstrating robustness under diverse seasonal and illumination changes.
Problem

Research questions and friction points this paper is trying to address.

Cross-Modal Place Recognition
Vision Foundation Models
Depth Image Representation
Geometry-Aware Overlap Miner
Robust Localization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Depth-Based Cross-Modal Place Recognition
Vision Foundation Models
Geometry-Aware Overlap Mining
Modality-Invariant Global Descriptors
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yungsoo Han
Department of Aerospace Engineering, Seoul National University, Seoul, Republic of Korea
Youngseok Jang
Youngseok Jang
Seoul National University
SLAMCollaborative SLAMPerception-aware path planning
S
Seungwon Roh
Department of Aerospace Engineering, Seoul National University, Seoul, Republic of Korea
Jeongyeon Seo
Jeongyeon Seo
M.S student at KAIST
LLMResponsible AI
H. Jin Kim
H. Jin Kim
Seoul National University, South Korea
roboticsdronesintelligent control