LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments

📅 2026-07-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of traditional world models that rely on pixel-level video generation, which are prone to visual distractions and inefficient representations. The authors propose LeapBot-WA, a novel predictive latent paradigm that models physical dynamics in a semantically aligned latent space via a Joint Embedding Predictive Architecture (JEPA), eschewing explicit visual synthesis. The approach integrates an Isotropic Semantic Autoencoder (ISAE) with an asymmetric Mixture-of-Transformers (MoT) and incorporates a diffusion prior to enable efficient inference and strong robustness. LeapBot-WA achieves state-of-the-art performance among predictive models on LIBERO and matches top generative models on RoboTwin 2.0—without requiring large-scale trajectory pretraining—while demonstrating exceptional zero-shot generalization and real-world transfer capabilities.
📝 Abstract
World Action Models (WAMs) have emerged as a powerful paradigm for embodied intelligence, yet the prevailing reliance on pixel-level video generation creates a fundamental bottleneck. Forcing models to reconstruct task-irrelevant visual details dissipates representational capacity and renders policies vulnerable to visual distractors. In this paper, we propose LeapBot-WA, which establishes a novel Predictive-Latent paradigm for WAMs by operationalizing the Joint-Embedding Predictive Architecture (JEPA) as a World-Anchor. Departing from the traditional reliance on visual synthesis, LeapBot-WA shifts the core of world modeling to Predictive Semantic Alignment, extracting abstract physical dynamics directly within a latent foundation space. To bridge the modality gap between non-Gaussian predictive features and diffusion priors, we introduce the Isotropic Semantic Autoencoder (ISAE), which reshapes the anchor's latent space into a diffusion-friendly manifold to prevent off-manifold drift. Furthermore, we design an Asymmetric Mixture-of-Transformers (MoT) architecture. During training, an Anchor Diffusion Transformer acts as a privileged dynamics expert to guide the Action Diffusion Transformer; at inference, this heavy dynamics branch is pruned, enabling zero-overhead execution. LeapBot-WA achieves state-of-the-art performance among predictive models on LIBERO and matches top-tier generative WAMs on RoboTwin 2.0 without requiring large-scale trajectory pre-training. It further demonstrates superior zero-shot robustness to unseen environments and successful real-world transfer, establishing a highly efficient and robust latent-centric paradigm for scalable robotic control. Code: https://github.com/LeapWM/leapbot-wa.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
pixel-level video generation
visual distractors
representational capacity
embodied intelligence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Predictive Latent Alignment
World-Anchor Action Models
Isotropic Semantic Autoencoder
Asymmetric Mixture-of-Transformers
Diffusion-Friendly Manifold
💼 Related Jobs
No related jobs found.
Pei Liu
Pei Liu
The Hong Kong University of Science and Technoly
End-to-end Autonomous DrivingLarge Language Models
N
Nan Zheng
Leapmotor
L
Lang Zhang
Leapmotor
D
Daojie Peng
The Hong Kong University of Science and Technology (Guangzhou)
Y
Yanan Zhang
Leapmotor
F
Feilong Kong
Southeast University
M
Mingyue Feng
Leapmotor
J
Jiachao Liu
Leapmotor
Y
Yaonong Wang
Leapmotor
Qifeng Chen
Qifeng Chen
HKUST
Computational PhotographyImage SynthesisGenerative AIAutonomous DrivingEmbodied AI
Jun Ma
Jun Ma
Assistant Professor, The Hong Kong University of Science and Technology
RoboticsAutonomous DrivingMotion Planning and ControlOptimization