JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling
This work addresses the high deployment cost of existing video-based generative world models and the suboptimal control performance caused by decoupling state prediction from action modeling in latent approaches. It introduces, for the first time, the Joint-Embedding Predictive Architecture (V-JEPA) to world modeling, constructing a shared predictor in the pretrained V-JEPA latent space that jointly learns state transitions and continuous action generation. A structured current–future joint objective preserves dense visual correspondences, enabling tight coupling between prediction and policy. The resulting model integrates seamlessly into vision–language–action (VLA) systems. Evaluated on LIBERO-Plus, it achieves a 79.2% success rate—the best among methods without large-scale robot pretraining—and reaches a new state-of-the-art overall performance of 86.3% with the π₀.₅ instantiation, while demonstrating strong generalization on RoboTwin 2.0 and real-world dual-arm tasks.