RoboEdit: Turning Human Manipulation Videos into Scalable Robot Experience

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决机器人手-物体交互数据收集成本高且特定于实体的问题,本文提出RoboEdit系统,通过将人类操作视频转换为适合机器人学习的数据,使用自动处理流程生成大规模训练数据集。
📝 Abstract
Collecting robot hand-object interaction data is costly and embodiment-specific, yet abundant human-object videos remain unusable for robot training. We present RoboEdit, a human-to-robot video editing suite that transforms human manipulation videos into action-consistent, physically plausible robot videos with aligned 3D hand states. To enable scalable supervision, we introduce RoboEdit-ADC, an automatic pipeline that reconstructs and retargets 3D interactions from RGB videos across embodiments. This pipeline generates RoboEdit-14M, a large-scale dataset of 174K aligned video pairs (14M frames) spanning seven robot embodiments, diverse scenes, and interaction types. The core editing engine, RoboEdit-Trans, employs cross-embodiment adaptation modules to preserve temporal coherence while adapting appearance and motion. It further integrates a 3D Robot-State Decoder to recover per-frame hand states for structured motion supervision. Experiments show that RoboEdit achieves state-of-the-art editing quality and supports downstream robot control policies in real-world manipulation tasks. Ultimately, the RoboEdit suite unlocks the vast potential of unlabeled human videos, providing scalable, high-fidelity visual and 3D motion supervision for generalizable robot learning.
Problem

Research questions and friction points this paper is trying to address.

robot hand-object interaction
human manipulation videos
embodiment-specific
video editing
scalable supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Human-to-robot video editing
Cross-embodiment adaptation
3D Robot-State Decoder
Scalable supervision