Monkey See, Can Monkey Do? A Benchmark for Evaluating Robot Skill Learning by Observation

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对机器人通过观察学习技能难以评估的问题,提出了一种名为RoboReel的统一基准,用于评估从人类视频中学习策略的模型。
📝 Abstract
Learning from Observation (LfO) is a fundamental robotic capability that replicates how humans and animals socially learn from each other. Beyond its biological parallels, this modality provides a practical solution for data scaling in sample-inefficient and data-starved domains like robotics. Recent work has demonstrated promising results in learning manipulation skills from human videos, yet progress in this area remains difficult to assess. Existing methods vary widely in assumptions, hardware choices, and environment setups making it difficult to draw meaningful comparisons and identify advances in the field. To address these challenges, we introduce RoboReel: a unified benchmark for evaluating models that learn policies from human videos. RoboReel consists of bundled real-world human demonstration videos, simulated robot trajectories, and evaluation environments on ten manipulation tasks. We develop four test suites to evaluate the models'performance on multiple axes, including the robustness to visual distractors and the ability to complete long-horizon tasks. Our benchmark covers learning-from-observation models from different categories, and studies the effectiveness of multiple representation choices in our benchmark evaluation that covers over seven state-of-the-art algorithms (including our VLA based variants) in the field of LfO. Finally, we present an analysis of the different types of algorithms showing that long-horizon tasks and tasks with low tolerances are still challenging for current models. Webpage: https://roboreel.github.io
Problem

Research questions and friction points this paper is trying to address.

Learning from Observation
Robotics
Benchmark Evaluation
Manipulation Tasks
Long-horizon Tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Learning from Observation
Benchmark
Manipulation Tasks
Visual Distractors
Long-horizon Tasks
💼 Related Jobs
No related jobs found.
Weiwei Gu
Weiwei Gu
Beijing University of Chemical Technology
Network Representation LearningNetwork Geometric LearningDeep Reinforcement Learning
A
Anmol Gupta
School of Augmented Intelligence, Arizona State University
A
Anant Sah
School of Augmented Intelligence, Arizona State University
R
Ryan Varghese
School of Augmented Intelligence, Arizona State University
L
Lalitha Shreya Vanam
School of Augmented Intelligence, Arizona State University
P
Prabhath Adireddi
School of Augmented Intelligence, Arizona State University
Peter Karkus
Peter Karkus
Research Scientist, NVIDIA Research
RoboticsArtificial IntelligenceMachine LearningAutonomous Vehicles
Nakul Gopalan
Nakul Gopalan
Assistant Professor Arizona State University
RoboticsNatural LanguageReinforcement Learning