SV-WAM: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决自动驾驶中全视角覆盖与计算效率的问题,SV-WAM通过共享生成模型进行动作学习,并采用可行驶区域合规正则化器提高安全性。
📝 Abstract
World models (WMs) have demonstrated strong potential for end-to-end autonomous driving by learning predictive representations of future scene dynamics. However, generating future videos during inference introduces substantial computational overhead, leading many recent driving WMs to adopt a single front camera as input for efficient deployment. This design restricts spatial coverage in safety-critical maneuvers such as lane changes, merges, and turns. To address this limitation, we propose SV-WAM, a surround-view world-action model (WAM) that preserves full six-camera observations while maintaining efficient inference. SV-WAM leverages future-video prediction as dense training supervision for action learning within a shared generative model, rather than as an inference-time output. At the core of this design is an action-centered causal mask that prevents action tokens from attending to future-video tokens during joint action-video denoising. Consequently, the video branch can be discarded at deployment, enabling efficient action-only planning. Furthermore, we introduce a differentiable drivable-area compliance regularizer that penalizes vehicle-footprint corners approaching or crossing drivable boundaries, improving planning safety and boundary awareness. Extensive experiments on the closed-loop NAVSIMv2 benchmark and the open-loop nuScenes benchmark demonstrate that SV-WAM achieves state-of-the-art planning performance with low inference latency and competitive zero-shot transfer capability.
Problem

Research questions and friction points this paper is trying to address.

end-to-end autonomous driving
computational overhead
spatial coverage
Innovation

Methods, ideas, or system contributions that make the work stand out.

surround-view world-action model
action-centered causal mask
drivable-area compliance regularizer
🔎 Similar Papers
J
Jinyang Wang
Institute of Automation, Chinese Academy of Sciences
S
Shiwei Li
Chongqing Changan Technology Co., Ltd.
J
Junjian Wang
Institute of Automation, Chinese Academy of Sciences
Zhiqiang Deng
Zhiqiang Deng
Chongqing Changan Technology Co., Ltd.
J
Jianbin Gao
Civil Aviation University of China
Y
Yihang Zhao
Institute of Automation, Chinese Academy of Sciences
L
Liu Liu
Chongqing Changan Technology Co., Ltd.
Y
Yongjia Zhao
Beihang University
J
Jinlong Chen
Guilin University of Electronic Technology
H
Huirui Xu
Institute of Automation, Chinese Academy of Sciences
Y
Yifeng Pan
Chongqing Changan Technology Co., Ltd.
Kangwei Liu
Kangwei Liu
Institute of Information Engineering, Chinese Academy of Sciences
Audio-driven Talking Face GenerationFacial Animation
Fan Ren
Fan Ren
Chongqing Changan Technology Co., Ltd.
J
Ji Tao
Chongqing Changan Technology Co., Ltd.
M
Minghao Yang
Institute of Automation, Chinese Academy of Sciences