DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出DECOWAM模型,通过分离相机自运动与基座和臂动作因素来改进移动操作中的未来观察与控制预测。
📝 Abstract
Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish camera ego-motion from base and arm actions. Here we introduce DECOWAM, a whole-body world-action model that separates these factors through dedicated conditional interfaces. DECOWAM freezes an adapted FastWAM backbone and trains residual adapters, an action-equivalent future bottleneck distilled from privileged observations, adversarially separated base and arm latents, and base-velocity conditioning for video prediction. We further introduce ARMDOG, a real-robot dataset that synchronizes video, whole-body state and action, and language. On a fixed replay protocol, DECOWAM improved both future-video and action prediction over FastWAM, reducing action MSE by 21.7% with 25.95M trainable adaptation parameters. Across 79 closed-loop trials per method, it achieved the highest observed whole-body coordination and base-displacement robustness among the compared systems, while task completion remained comparable to the strongest baseline. These results show that embodiment-aware factorization can support parameter-efficient joint visual prediction and whole-body control under moving viewpoints.
Problem

Research questions and friction points this paper is trying to address.

mobile manipulation
world-action model
camera ego-motion
base and arm actions
Innovation

Methods, ideas, or system contributions that make the work stand out.

DECOWAM
residual adapters
action-equivalent future bottleneck
adversarially separated base and arm latents
base-velocity conditioning
💼 Related Jobs
No related jobs found.
S
Siyuan Ma
Tsinghua University, Beijing, China
B
Boshi Zhang
Tsinghua University, Beijing, China
Y
Yutian Zhang
Shanghai Artificial Intelligence Laboratory, Shanghai, China
Q
Qinglian Wu
Harbin Institute of Technology, Harbin, China
J
Jiaqi Zhai
Hangzhou Yunshenchu Technology Co., Ltd. (DEEP Robotics), Hangzhou, China
D
Dong Wei
Hangzhou Yunshenchu Technology Co., Ltd. (DEEP Robotics), Hangzhou, China
Qiaojun Yu
Qiaojun Yu
Shanghai Jiao Tong University, Shanghai AI Lab
robotic learning3D visionvla