Rethinking Language's Role in Efficient VLA for Autonomous Vehicles: Toward Smarter, Trustworthy Driving

📅 2026-08-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了在自动驾驶中何时何地使用语言以降低成本,并提出语言残差分类法来组织不同方法,分析其在推理时的效率和效果。
📝 Abstract
Vision-Language-Action (VLA) models are reshaping autonomous driving (AD) by unifying perception, reasoning, and control through language, enabling semantic grounding, interpretable decisions, and better long-tail generalization. But language is expensive onboard: latency and memory budgets are tight, and autoregressive decoding is inherently sequential. This work reframes the central question as when and where language should act at inference, since inference cost recurs at every deployed frame while training cost is paid once. We introduce the Language Residue taxonomy to organize methods by their inference-time use of language: train-time-only supervision (L1), latent non-textual reasoning (L2), conditional invocation (L3), and full per-frame generation (L4). We review representative methods and tag each across five deployment axes (latency, parameters, memory, FLOPs, tokens), analyzing them on major open- and closed-loop driving benchmarks (e.g., nuScenes, NAVSIM, Bench2Drive). We further trace how efficient methods from NLP/LLM are adapted in AD, identifying the constraints and motivations driving these adaptations. A continuously updated repository will be available at Github.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
autonomous driving
inference cost
language utilization
efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
Language Residue
autonomous driving
efficiency
inference cost
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
T
Tongfei Guo
Department of Electrical and Computer Engineering, Northeastern University, Boston, MA, USA
Lili Su
Lili Su
Assistant Professor, Northeastern University
Distributed learningmachine learningFault/adversary-tolerant computingperformance evaluation