The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入模仿游戏及配套数据集IG-10K,解决机器人从人类视频学习时意图理解不足的问题,评估了多种模型在不同模仿难度下的表现。
📝 Abstract
Humans imitate at the level of intent: given a demonstration, we infer its goal and carry it out with whatever tools, objects, and layouts are at hand. Current robot policies instead learn observation-to-action mappings from visual inputs and language instructions, without explicitly inferring the demonstrated task. Learning from human video thus remains largely trajectory-level: models can replay motions in near-identical scenes, but still struggle to imitate what the demonstrator intends rather than merely what they do. We introduce The Imitator Game, a four-level benchmark (L0-L3) that progressively widens the gap between the human demonstration and the robot's own scene, isolating where trajectory replay ceases to suffice and task understanding becomes necessary. We pair it with IG-10K, the largest environment-aligned paired human-robot dataset to date and the only one instantiated across all four levels in both real and simulated settings (20,000+ paired episodes, 50+ tasks, 6 domains), and Imitator Arena, an open platform for blind A/B human evaluation. Across nine state-of-the-art models, performance is stable from L0 to L2 but collapses at L3, identifying functional substitution - achieving the same intent through a different object affordance - as the decisive barrier to intent-level imitation. Human-video-conditioned models outperform caption-conditioned ones, yet every model falls below 13% zero-shot success on unseen tasks; fine-tuning IG-10K-pretrained models with only $10$ paired human-robot demonstrations yields large gains that grow with pretraining scale. The project website and access to Imitator Arena are available at https://imitator-game.github.io.
Problem

Research questions and friction points this paper is trying to address.

imitation learning
intent understanding
trajectory replay
functional substitution
human-robot interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Intent-level Imitation
Imitator Game
IG-10K Dataset
Functional Substitution
💼 Related Jobs
No related jobs found.
X
Xunzhe Zhou
The University of Hong Kong
Y
Yiyang Cai
Fudan University
F
Fengyi Wang
Fudan University
R
Ran Ju
The University of Hong Kong
H
Hanxiang Ren
Zhejiang University
R
Ruizhe Liu
The University of Hong Kong
Y
Yu Zhang
The University of Hong Kong
Qian Luo
Qian Luo
Assistant Research Professor, The George Washington University
Feng Chen
Feng Chen
Southwest University, Chongqing, China
signal processingdistributed estimationdistributed signal processingPoint Cloud
P
Pei Zhou
The University of Hong Kong
Y
Yi Ma
The University of Hong Kong
Yanchao Yang
Yanchao Yang
Assistant Professor, HKU; Stanford University; UCLA
Embodied AIComputer VisionMachine Learning