DA-WAM: Decision-Aligned Future Latents for Driving World Models

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决自动驾驶中场景预测与决策分离的问题,提出DA-WAM框架,通过统一表示学习、行动条件未来建模和轨迹评分来直接指导决策。
📝 Abstract
Anticipating how scenes evolve under ego actions is fundamental to safe autonomous driving, yet the full potential of world models for decision-making remains unrealized. The critical challenge lies in ensuring that future modeling is not merely predictive, but decision-informative: the predicted future must directly shape which trajectory is selected. Existing approaches decouple future representation learning from planning optimization, or share predicted states across trajectory candidates, thereby diluting the action-specific consequences that ought to guide selection. To bridge this gap, we propose DA-WAM, a framework that unifies predictive representation learning, action-conditioned future modeling, and trajectory scoring under a single decision-making objective. DA-WAM maintains predictive supervision throughout planner optimization via an online encoder and a stable momentum target, allowing future representations to co-evolve with the driving task. An action-conditioned predictor generates a distinct future latent state per trajectory candidate, which is then evaluated by a future-latent-conditioned factorized scorer. For the expert-matched trajectory, the predicted future latent is supervised by the observed future representation, while safety-critical hard negatives provide additional supervision near planning boundaries. Extensive experiments on NAVSIM-v1 and NAVSIM-v2 demonstrate state-of-the-art performance, while ablations and diagnostic analyses validate the key components.
Problem

Research questions and friction points this paper is trying to address.

autonomous driving
world models
decision-making
future modeling
trajectory selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decision-Aligned
Action-Conditioned
Future Latents
Trajectory Scoring
Predictive Supervision
R
Ruiguo Zhong
The Hong Kong University of Science and Technology (Guangzhou)
B
Benshan Ma
The Hong Kong University of Science and Technology (Guangzhou)
X
Xiaolong Chen
The Hong Kong University of Science and Technology (Guangzhou)
L
Lang Zhang
Leapmotor
M
Mingyue Feng
Leapmotor
Y
Yaonong Wang
Leapmotor
Pei Liu
Pei Liu
The Hong Kong University of Science and Technoly
End-to-end Autonomous DrivingLarge Language Models
J
Jun Ma
The Hong Kong University of Science and Technology (Guangzhou)