VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models
This work addresses the challenge of accurate localization from natural language descriptions to 3D point cloud maps by proposing VLM-Loc, a novel approach that leverages vision-language models (VLMs) for cross-modal alignment between text and point clouds. By transforming point clouds into bird’s-eye-view images and structured scene graphs, the method jointly encodes geometric and semantic information. A partial node matching mechanism is introduced to enable interpretable spatial reasoning. Evaluated on the newly constructed CityLoc benchmark, VLM-Loc significantly outperforms existing methods, achieving state-of-the-art performance in both localization accuracy and robustness, while simultaneously enhancing the model’s spatial reasoning capabilities and decision interpretability.