WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving

📅 2026-08-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对自动驾驶规划需求,提出WA-JEPA模型,通过混合未来掩码预训练和条件流匹配方法改进了V-JEPA的未来预测能力。
📝 Abstract
Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent feature prediction. However, V-JEPA is built around random-mask completion and deterministic regression, making it fundamentally ill-suited for autonomous driving planning that demands future-directed prediction tightly coupled with action. To address this, we rethink the V-JEPA paradigm and present WA-JEPA, a V-JEPA-native world-action model designed for autonomous driving planning. Instead of random spatiotemporal masking, WA-JEPA employs hybrid future-masked pre-training, where the model infers future latents from observed context. Departing from deterministic regression, we recast future prediction as conditional flow matching over latent futures, which substantially improves the model's ability to generate plausible future latents for downstream planning. Finally, a joint future-action predictor is proposed to denoise future scene tokens and ego trajectories together in a unified spatiotemporal latent space, allowing action supervision to directly shape planning-relevant world representations. Pre-trained on nuPlan videos and fine-tuned on NAVSIM, WA-JEPA reaches 91.7 EPDMS on NAVSIM-v2, surpassing the strongest end-to-end and world-action baselines by 1.6 and 1.3 EPDMS, and, without HUGSIM-specific fine-tuning, attains the best HD-Score of 0.4462 on the closed-loop HUGSIM benchmark under the same evaluation protocol. These results validate V-JEPA-native world-action modeling as a powerful and scalable paradigm for autonomous driving planning. Code is available at https://github.com/AFARI-Research/WA-JEPA.
Problem

Research questions and friction points this paper is trying to address.

V-JEPA
autonomous driving planning
future prediction
spatiotemporal representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

hybrid future-masked pre-training
conditional flow matching
joint future-action predictor
🔎 Similar Papers
No similar papers found.
X
Xinlin Wang
Afari Intelligent Drive
Y
Yujiao Xiang
Afari Intelligent Drive; University of Electronic Science and Technology of China
Y
Yuheng Zhou
Afari Intelligent Drive; Southeast University
J
Jingqi Wang
Afari Intelligent Drive
M
Minqing Huang
Afari Intelligent Drive
J
Jiajie Huang
Afari Intelligent Drive; Beijing University of Posts and Telecommunications
D
Dongxu Wei
Afari Intelligent Drive
T
Tingguang Zhou
Afari Intelligent Drive
X
Xiyang Wang
Afari Intelligent Drive
Gong Chen
Gong Chen
Nanjing University
Magnetic imaging
Z
Zhi Xu
Afari Intelligent Drive
F
Feiyang Tan
Afari Intelligent Drive
H
Hangning Zhou
Afari Intelligent Drive
M
Mu Yang
Afari Intelligent Drive