Continuous Actions from Discrete Minds: Latent-Aligned Planning for End-to-End Autonomous Driving

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入LaPla框架,利用潜变量对齐规划来解决视觉-语言模型离散推理与自动驾驶连续物理约束之间的差距,提高驾驶平滑性和成功率。
📝 Abstract
Bridging the gap between the discrete reasoning of Vision-Language Models and the continuous, physics-constrained nature of autonomous driving remains a significant challenge. In this work, we introduce LaPla, a unified Vision-Language-Action (VLA) framework featuring latent-aligned planning to seamlessly ground semantic understanding in precise motion execution. We first design an action tokenizer based on a residual vector-quantized variational autoencoder (VQ-VAE), capturing vehicle kinematics and encoding trajectory features into a structured latent space. Rather than discrete codebook lookups that inevitably introduce quantization errors, LaPla repurposes this representation as a physical prior to bridge the modality gap between high-dimensional semantics and the raw action space. Specifically, given multimodal inputs integrating multi-view images, historical actions, and textual instructions, LaPla incorporates concurrent action queries to causally attend to the multimodal context in a single forward pass, projecting hidden states directly into the pretrained VQ-VAE latent space. The frozen decoder then translates these continuous latents into actions, effectively eliminating quantization errors and ensuring physically plausible trajectories while bypassing time-consuming autoregressive generation. Extensive experiments on the nuScenes benchmark demonstrate that LaPla achieves competitive open-loop performance, reducing long-horizon L2 error by 15.52% compared to state-of-the-art VLA methods. Closed-loop evaluations on the NVIDIA AlpaSim simulator further confirm its superior capability in ensuring smooth driving progress, improving the success rate by 33.34 percentage points with significantly reduced inference latency.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Models
Autonomous Driving
Continuous Actions
Discrete Reasoning
Physics-Constrained
Innovation

Methods, ideas, or system contributions that make the work stand out.

latent-aligned planning
VQ-VAE
multimodal context
continuous latents
autoregressive generation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
R
Ruoyu Yao
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou 511453, China
Y
Yusen Xie
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou 511453, China
Q
Qingzhao Liu
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou 511453, China
Pei Liu
Pei Liu
The Hong Kong University of Science and Technoly
End-to-end Autonomous DrivingLarge Language Models
Z
Zewei Yang
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou 511453, China
Y
Yipeng Zhu
Central Media Technology Institute, Huawei
X
Xiaolong Wang
Central Media Technology Institute, Huawei
J
Jun Ma
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou 511453, China