WholeBodyWAM: Learning Whole-Body World Action Models with Scalable Motion Priors

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出WholeBodyWAM模型,利用大规模异源运动数据作为先验,通过预训练和机器人后训练相结合的方法,解决人形机器人全身动态建模问题。
📝 Abstract
Humanoid whole-body manipulation requires coordinated whole-body dynamics, yet large-scale trajectories from a target robot are expensive to collect and difficult to scale. In contrast, whole-body motion from human and humanoid sources is abundantly available, although such data cannot be directly used as embodiment-specific robot actions. This work asks whether these scalable motion resources can instead provide a transferable predictive prior for humanoid world-action modeling. We introduce WholeBodyWAM, a humanoid world-action model that learns whole-body dynamics from large-scale heterogeneous motion before target-robot training. We curate UniMotion-4K, a motion corpus spanning more than 4K hours from human videos, native 3D motion datasets, and heterogeneous humanoid platforms, and canonicalize these diverse sources into a unified motion space. A language-conditioned Motion Expert is then pretrained to predict future whole-body motion without target-robot action supervision. During robot post-training, the pretrained Motion Expert is integrated with Video and Action Experts through asymmetric Mixture-of-Transformers (MoT) attention, enabling predictive scene dynamics and whole-body motion to jointly inform embodiment-specific action generation. Experiments show that WholeBodyWAM consistently benefits from increased motion-pretraining scale, improves future-motion prediction and downstream task performance, and transfers effectively to real-world humanoid manipulation. Moreover, the pretrained motion prior substantially improves data efficiency under limited target-robot demonstrations.
Problem

Research questions and friction points this paper is trying to address.

Whole-Body Manipulation
Scalable Motion Priors
Humanoid World-Action Modeling
Motion Data
Embodiment-Specific Actions
Innovation

Methods, ideas, or system contributions that make the work stand out.

WholeBodyWAM
Scalable Motion Priors
Asymmetric Mixture-of-Transformers (MoT)
Unified Motion Space
Data Efficiency
🔎 Similar Papers
No similar papers found.
Bowei Zhang
Bowei Zhang
Peking University
Q
Qiyao Zhang
Beijing Innovation Center of Humanoid Robotics
Shuanghao Bai
Shuanghao Bai
Xi'an Jiao Tong University Phd student
Vision Language ModelsDomain AdaptationDomain GeneralizationRobotic Manipulation
X
Xinhua Wang
Beijing Innovation Center of Humanoid Robotics
M
Meng Li
Beijing Innovation Center of Humanoid Robotics
Yilei Wang
Yilei Wang
Alibaba Cloud
L
Leiwang Zhang
Tsinghua University
J
Jian Tang
Beijing Innovation Center of Humanoid Robotics
L
Lu Zhou
Nankai University
L
Lei Sun
Nankai University
Zhengping Che
Zhengping Che
X-Humanoid
Embodied AIDeep Learning