FetchMan: Learning Visual Humanoid Loco-Manipulation Policies from Simulated Experiences

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过模拟数据学习并转移到现实世界的方法,解决了机器人在新场景和物体上进行视觉操作与移动的难题,使用强化学习提高了性能。
📝 Abstract
Visual loco-manipulation policies that can generalize to novel scenes and objects have long been a goal of robotics research. However, today's data-hungry algorithms make collecting sufficient demonstrations a struggle for tabletop manipulation, and even more so for humanoids that must also walk and balance. Learning from simulated data and transferring that behavior to the real world, as is commonly done in locomotion, sidesteps this struggle, so we replicate that recipe for loco-manipulation. In doing so, we find that cloning synthetic demonstrations results in a low performance ceiling no matter the amount of training data. Reinforcement learning breaks through it, and refining the cloned policy with Flow-GRPO on a single sparse reward yields performance that synthetic behavior cloning cannot match. Together, these stages form our end-to-end sim-to-real pipeline spanning more than 150,000 scenes, which we use to train FetchMan. We evaluate it on FetchMan-Bench, a simulation benchmark we release, and deploy it zero-shot on a real Unitree G1, where our single-object reach-and-pick policy walks to and grasps a target across unseen scenes at 73.3% success. Finally, we extend this recipe to multi-object training, a first step toward loco-manipulation generalist policies at this data scale.
Problem

Research questions and friction points this paper is trying to address.

Visual Loco-Manipulation
Simulated Experiences
Generalization
Humanoid Robots
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual loco-manipulation
sim-to-real transfer
reinforcement learning
Flow-GRPO
🔎 Similar Papers
2024-07-16Neural Information Processing SystemsCitations: 16
💼 Related Jobs
No related jobs found.
Omar Rayyan
Omar Rayyan
UCLA
RoboticsMachine Learning
Z
Zhi Li
University of California, Los Angeles
Max Argus
Max Argus
University of Freiburg
CV | ML | Robotics
Y
Yuxin Jiang
University of California, Los Angeles
C
Chang Yu
University of California, Los Angeles
Chenfanfu Jiang
Chenfanfu Jiang
Professor, UCLA
Computer GraphicsComputer VisionEmbodied AIRobotics
Y
Yuchen Cui
University of California, Los Angeles