WLA$^3$: World Latent Action Modeling for Semantics, Dynamics, and Kinematics

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出WLA³模型,通过学习世界状态变化来解决多源数据下动作监督不足的问题,利用人类第一视角视频和机器人轨迹提高动作识别准确性。
📝 Abstract
Scaling generalist policy models with heterogeneous data is limited by the lack of unified, low-noise action supervision. Human egocentric videos are abundant, but only a small fraction comes with high-quality hand-action labels. Observed world transitions offer a common source of action-related supervision across data sources. We introduce WLA$^3$ (World Latent Action Modeling for Semantics, Dynamics, and Kinematics), a unified generalist policy model framework built around representations learned by a World Latent Action Model (WLAM). WLAM first learns how multimodal world states change over a local interval, encoding synchronized camera views and available embodiment-state changes into a compact local latent action and a richer transition feature. Reconstruction from partial modalities and consistency across overlapping windows encourage robust transition representations. WLA$^3$ reuses them across semantics, dynamics, and kinematics: local latent actions support action-sensitive physical-dynamics modeling, segment-level features directly supervise the VLM through a Semantic Latent Aggregate (SLA), and an action expert jointly predicts latent actions together with embodiment-specific robot controls. Human videos provide scalable transition supervision, while robot trajectories ground the shared representation in executable native controls. On LARYBench, the final 32D latent action reaches 67.89\% average classification accuracy. WLA$^3$ achieves 81.9% average success across six real-robot tasks versus 66.2% for $π_{0.5}$. Performance improves as generalist policy model mid-training data scales, and human videos support human-to-robot transfer. Project page can be found at https://wla-3.github.io/.
Problem

Research questions and friction points this paper is trying to address.

generalist policy models
heterogeneous data
action supervision
human egocentric videos
hand-action labels
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Latent Action Model
Unified Generalist Policy Model
Latent Action
Semantic Latent Aggregate
Transition Supervision
💼 Related Jobs
No related jobs found.
Peidong Liu
Peidong Liu
Westlake University
3D computer visionRobotics
Z
Zhiyuan Xiang
Joy Future Academy, JD Group
M
Mingyang Li
Joy Future Academy, JD Group
W
Wenhao Li
Joy Future Academy, JD Group
J
Jiale Zhang
Joy Future Academy, JD Group
J
Jiahao Sun
Joy Future Academy, JD Group
J
Jiawei Li
Joy Future Academy, JD Group