Pre-training Visual Dexterity in Simulation

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the scarcity of dexterous manipulation data and the high cost of real-world collection by proposing the SPD framework. The method leverages large-scale VR-simulated teleoperation data to pretrain a causal Transformer, followed by behavioral cloning fine-tuning with minimal real-world demonstrations, thereby validating simulated data as an effective pretraining source for physical manipulation. Experiments demonstrate that efficient transfer is achievable with only one to two hours of real-robot data, yielding performance significantly superior to behavioral cloning policies trained from scratch. These findings establish a novel paradigm for acquiring high-quality dexterous manipulation models at reduced cost, effectively bridging the sim-to-real gap through strategic pretraining and lightweight adaptation.
📝 Abstract
Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this progress has largely been driven by datasets and embodiments built around simple parallel-jaw grippers. Dexterous, multi-fingered hands remain comparatively data-starved because real teleoperation is costly to scale, while human hand video is off-embodiment and requires lossy pose estimation and retargeting. We introduce Simulation Pre-training for Dexterity (SPD), a pre-training framework for dexterous manipulation that uses data entirely collected in simulation. In SPD, humans manipulate virtual objects inside a VR headset, enabling on-embodiment trajectories and robot-free collection. With the help of five operators, we collect 75 hours of multi-task dexterous manipulation over one week, and use it to pre-train a causal transformer on a sequence modeling objective. We study the benefits of simulation pre-training on real-world tasks by fine-tuning on 1-2 hours of physical demonstrations on a 56-DoF bimanual dexterous setup. We find that our approach outperforms training behavior cloning policies from scratch, showing that simulation teleoperation is a viable pre-training source for real-world dexterous manipulation. We perform ablation studies, measuring the benefits of history conditioning and short action chunks for reactive control.
Problem

Research questions and friction points this paper is trying to address.

Dexterous Manipulation
Data Scarcity
Pre-training
Cross-embodiment
Teleoperation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Simulation Pre-training
Dexterous Manipulation
VR Teleoperation
Causal Transformer
Sim-to-Real Transfer
🔎 Similar Papers
No similar papers found.