Drive-HWM: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决自动驾驶中长期预测与即时决策难以兼顾的问题,提出Drive-HWM框架,通过慢-快层级模型在不同时间尺度上进行未来场景预测和动作生成。
📝 Abstract
World models offer a promising paradigm for autonomous driving by predicting how traffic scenes may evolve and using such predictions to support action generation. However, existing approaches either separate future prediction from action generation or jointly predict them at the same temporal scale, making it difficult to simultaneously achieve long-horizon anticipation and responsive, observation-grounded decision making. We present Drive-HWM, a hierarchical slow--fast world modeling framework that organizes future representation prediction and action generation at complementary temporal scales. The slow world model predicts multi-step future representations to capture extended scene evolution. To explicitly model the abundant motion dynamics in driving environments, we introduce Dynamic-Aware Latents learned through optical-flow prediction. Guided by these future representations, the fast model uses a lightweight multimodal backbone and an autoregressive expert to jointly predict the next frame and the immediate action from the latest observation. Next-frame prediction encourages the fast model to capture imminent scene evolution, while one-step action generation allows decisions to be continuously updated as new observations arrive. Extensive experiments on NAVSIM v1 and v2 demonstrate the strong driving performance of Drive-HWM. Comprehensive ablation studies further validate the effectiveness of the hierarchical slow--fast design, dynamics-aware future representations, and joint next-frame and action prediction.
Problem

Research questions and friction points this paper is trying to address.

autonomous driving
world models
long-horizon anticipation
responsive decision making
temporal scales
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical World Models
Dynamic-Aware Latents
Autonomous Driving
Slow-Fast Framework
Next-Frame Prediction
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhaoxin Fan
School of Artificial Intelligence, Beihang University, Beijing, China
T
Tianbao Zhang
Shanghai Jiao Tong University, Shanghai, China and Dim12 AI
W
Wenjun Wu
School of Artificial Intelligence, Beihang University, Beijing, China
X
Xiaofeng Wang
GigaAI
Yeying Jin
Yeying Jin
Tencent | National University of Singapore
Computer VisionAIGCGenAIMLLMVLM
J
Jian Zhao
TeleAI
Z
Zheng Zhu
GigaAI
S
Shuicheng Yan
National University of Singapore, Singapore