Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Geo-VLA框架,通过学习几何感知视觉表示来增强VLA模型,以解决复杂驾驶环境中图像表示不足的问题。
📝 Abstract
Vision-language-action (VLA) models have advanced end-to-end autonomous driving by leveraging foundation models for semantic reasoning and long-tail generalization. However, their planning performance remains limited in complex driving environments because image-only representations inadequately capture planning-relevant road geometry and topology. In this paper, we propose Geo-VLA, a plug-and-play framework that enhances VLA models by learning geometry-aware visual representations. During training, Geo-VLA internalizes geometric map semantics to strengthen road-structure representations, while requiring no HD maps or additional lane information during inference. To support this approach, we introduce Geo-QA, a geometry-focused question-answering dataset that injects road geometry into vision-language representations through contrastive learning and instruction tuning. Experiments on NAVSIM v1 demonstrate that Geo-VLA consistently improves VLA planners with distinct action-generation architectures, achieving 92.1 PDMS and establishing a new state-of-the-art among single-camera VLA planners.
Problem

Research questions and friction points this paper is trying to address.

Vision-language-action
planning performance
complex driving environments
road geometry
topology
Innovation

Methods, ideas, or system contributions that make the work stand out.

geometry-aware visual representations
internalization of geometric map semantics
Geo-QA dataset
🔎 Similar Papers