4D-WAM: 4D Consistent World Modeling for Autonomous Driving

πŸ“… 2026-08-10
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing world-action models, trained solely on 2D videos, struggle to capture the true 4D structure of dynamic scenes, often yielding predictions that appear visually plausible yet suffer from spatiotemporal inconsistencies that mislead downstream planning. To address this, this work proposes integrating 4D supervisory signals derived from geometric foundation models, introducing a 4D consistency loss, and employing a decision-aware temporal step sampling strategy that intensifies supervision during the high-noise early stages of training. The approach incurs no additional inference overhead while enabling physically consistent modeling of 4D scene evolution and improved trajectory planning. It achieves state-of-the-art performance on both NAVSIM-v1 and NAVSIM-v2 benchmarks.
πŸ“ Abstract
Emerging World-Action Models (WAMs) have demonstrated promising performance in autonomous driving by jointly modeling future driving scene evolution and trajectory planning. However, existing WAMs are typically trained with video data, which is only 2D projections of the underlying 4D driving scene. Consequently, WAMs fail to understand and capture the structure of 4D scenes and thus generate visually plausible yet 4D inconsistent future predictions that mislead downstream planning. To alleviate this issue, we present 4D-WAM, a model that leverages geometric foundation models for training-time supervision to enable 4D consistent world modeling. Specifically, we feed WAM-predicted future frames into a geometric foundation model, and use 4D-aware responses to define a 4D consistency loss. This loss encourages the model to understand, represent, and predict physically consistent 4D scenes during training, without additional inference cost. Moreover, we identify an early-decision phenomenon in WAMs and propose a decision-oriented timestep sampling strategy that emphasizes supervision at early, high-noise stages, where driving decisions are primarily formed. By propagating 4D supervision to this critical decision-formation phase, the proposed strategy further improves trajectory planning. Extensive experiments demonstrate that 4D-WAM effectively models 4D consistent scene evolution and achieves state-of-the-art performance on challenging NAVSIM-v1 and NAVSIM-v2 benchmarks.
Problem

Research questions and friction points this paper is trying to address.

4D consistency
world modeling
autonomous driving
future prediction
trajectory planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

4D consistency
World-Action Models
geometric foundation models
decision-oriented sampling
autonomous driving
πŸ’Ό Related Jobs
No related jobs found.
J
Jiacheng Fu
University of Science and Technology of China
Y
Yibo Yuan
Xi’an Jiaotong University
M
Meng Tian
Yinwang Intelligent Technology Co., Ltd.
Y
Yue Li
Yinwang Intelligent Technology Co., Ltd.
Jiangtong Zhu
Jiangtong Zhu
XJTU
Jianhua Han
Jianhua Han
2030 Research, YinWang, Huawei
Vision Language ModelFoundation ModelVLA
Yueyi Zhang
Yueyi Zhang
Miromind, Previously University of Science and Technology of China
Structured lightDepth SensingEvent CameraMedical Imaging
Jianwu Fang
Jianwu Fang
Xi'an Jiaotong University
Scene understandingSafe driving perception and planning
J
Jianru Xue
Xi’an Jiaotong University
H
Hang Xu
Yinwang Intelligent Technology Co., Ltd.
Zhiwei Xiong
Zhiwei Xiong
University of Science and Technology of China
computational photographybiomedical image analysis