Motus2: A Self-Evolving General World Model for Dexterous Manipulation

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Motus2,通过模型扩展和数据扩展方法解决灵巧操作中的感知、预测、行动、评估和改进问题。
📝 Abstract
General embodied agents should perceive, predict, act, evaluate, and improve within a unified system. World models have shown great promise in building such agents, yet existing models typically append an action output head to a world simulator, without coupling them into a closed decision-and-learning loop for policy improvement. We present Motus2, a self-evolving general world model for dexterous manipulation. Motus2 advances world modeling through model scaling and data scaling. For model scaling, a single model with shared weights exposes three control interfaces: a policy (world-action model), a simulator (action-conditioned world model), and an evaluator (value model). The policy proposes candidate action chunks, the simulator predicts their visual consequences, and the evaluator assesses the predicted outcomes. Their coupling forms a closed decision-and-learning loop for policy improvement. This formulation uses curated expert demonstrations for action learning, while failed and suboptimal interactions provide valuable evidence for dynamics modeling and value learning. For data scaling, Motus2 progresses from large-scale monocular egocentric data to synchronized stereo egocentric data, followed by robot-domain adaptation with robot trajectories and supplementary human-robot alignment data. Motus2 further studies global-autoregressive and hybrid-memory extensions of its sliding-window context, adds tactile feedback for contact-aware control, and is instantiated on a fully biomimetic platform with stereo vision, dual arms, dual dexterous hands, and tactile sensing. Together, egocentric data scaling and closed-loop general world model scaling provide a general path toward self-evolving dexterous manipulation.
Problem

Research questions and friction points this paper is trying to address.

world model
dexterous manipulation
policy improvement
embodied agent
closed decision-and-learning loop
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-evolving general world model
model scaling and data scaling
closed decision-and-learning loop
egocentric data scaling
tactile feedback
💼 Related Jobs
No related jobs found.
H
Hongzhe Bi
GensPI, Tsinghua University
Z
Zihao Zhou
GensPI, BUAA
Y
Yihang Tang
GensPI, BIT
J
Jingrui Pang
GensPI, Tsinghua University
S
Shuhe Huang
GensPI, Tsinghua University
H
Haitian Liu
GensPI, Tsinghua University
R
Runqing Wang
GensPI, BIT
S
Shuai Huang
GensPI
Y
Yichen Wang
GensPI
Yiming Cheng
Yiming Cheng
Tsinghua University
machine learningnetwork systemsdata miningrecommendation systems
Ruowen Zhao
Ruowen Zhao
Tsinghua University
3D VisionGenerative Model
Z
Zhenghua Li
Tsinghua University
Hengkai Tan
Hengkai Tan
Tsinghua University
Reinforcement LearningRobot LearningEmbodied AIDeep Generative Models
X
Xiaolong Liu
GensPI
J
Jinhui Wan
GensPI
J
Jiabao Liu
GensPI
Min Zhao
Min Zhao
Tsinghua University
Generative ModelsVision Generation
Fan Bao
Fan Bao
ShengShu
machine learning
Jun Zhu
Jun Zhu
Professor of Computer Science, Tsinghua University
Machine LearningBayesian MethodsDeep Generative ModelsAdversarial RobustnessReinforcement Learning