GRAVA: Grounded Reasoning-to-Action Representation and Learning for Autonomous Driving

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
针对自动驾驶中视觉-语言-动作模型推理与行为脱节问题,提出GRAVA框架,通过统一接地、推理和行动生成流程,增强物理场景证据与执行行为的关联。
📝 Abstract
Driving vision-language-action (VLA) models increasingly reason before acting, but their intermediate reasoning is often weakly grounded in physical scene evidence and loosely connected to executable behavior. We present GRAVA, a framework built around Grounded Reasoning-to-Action (GRA), which unifies grounding, reasoning, and action generation in a single autoregressive stream. GRA links action-relevant language references to 2D visual regions and ego-centric physical states, organizes object interactions and decisions in a trajectory-anchored typed graph, and serializes this structure into grounded reasoning. A single VLM generates this reasoning followed by a compact Executable Planner action that is deterministically decoded into a continuous trajectory. We further introduce an agentic GRA data construction pipeline that combines forward scene grounding with backward trajectory anchoring, and use it to build GR-NavSim with 2.2M grounded question-answer pairs and 70K GRA reasoning traces. A progressive training strategy develops grounded cognition through pre-training, establishes the reasoning-to-action interface through imitation, and improves driving behavior through reinforcement learning and exploration. Using about 60% of the available human driving demonstrations for action supervision, GRAVA-8B achieves state-of-the-art performance among purely autoregressive driving models on the full NAVSIM benchmark. On an internal long-tail benchmark, full GRA improves key-object compliance and Closed-loop Driving Score by 19.3% and 20.5% over action-only prediction, respectively. These results show the benefit of preserving action-relevant physical evidence from grounded reasoning through executable action generation.
Problem

Research questions and friction points this paper is trying to address.

Grounded Reasoning
Action Generation
Visual-Language-Action Models
Autonomous Driving
Physical Evidence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Grounded Reasoning-to-Action
autoregressive stream
trajectory-anchored typed graph
Executable Planner action
agentic GRA data construction
💼 Related Jobs
No related jobs found.
X
Xiao Liu
National Engineering Research Center of Electric Vehicles, Beijing Institute of Technology, Beijing 100081, China; Shenzhen Automotive Research Institute, Beijing Institute of Technology, Shenzhen 518118, China; Shenzhen Jiguangzhijie Technology Co., Ltd., Shenzhen 518118, China
Haoyu Li
Haoyu Li
National Institute of Informatics
speech processing
J
Jianghao Leng
National Engineering Research Center of Electric Vehicles, Beijing Institute of Technology, Beijing 100081, China; Shenzhen Automotive Research Institute, Beijing Institute of Technology, Shenzhen 518118, China; Shenzhen Jiguangzhijie Technology Co., Ltd., Shenzhen 518118, China
L
Lin Wang
School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 639798
Chao Sun
Chao Sun
Western Digital Research & Stanford Univ.
Machine LearningArtificial IntelligenceStorage architectureSSD controllerStorage Class Memory