🤖 AI Summary
This study addresses the challenges of data scarcity in dexterous manipulation and cross-embodiment generalization by proposing a unified learning framework. By constructing the OmniShare dataset and introducing the JAAS universal action representation, combined with SE(3) pose encoding and domain-adversarial learning, this approach effectively decouples and integrates human-robot demonstration data. Leveraging a Vision-Language-Action model, the proposed method achieves zero-shot human-to-robot skill transfer. Consequently, it significantly enhances cross-embodiment generalization capabilities and few-shot adaptation efficiency, establishing a novel paradigm for universal manipulation in embodied intelligence.
📝 Abstract
Dexterous manipulation is a fundamental capability for embodied intelligence, but scaling it remains difficult because robot demonstrations are expensive to collect and action spaces vary across embodiments. Policies trained on heterogeneous data can also entangle task-relevant visual cues with embodiment-specific appearance, limiting cross-embodiment generalization. We present AdvDex, a unified Vision-Language-Action framework for learning dexterous manipulation from human and robot demonstrations. First, we introduce OmniShare, a large-scale multimodal dataset of human manipulation demonstrations that provides high-quality kinematic supervision and tactile measurements while reducing reliance on robot teleoperation. Second, we propose the Joint-Aligned Action Space (JAAS), a canonical action representation comprising an $\mathrm{SE}(3)$ wrist pose and 15 finger joints, thereby functionally aligning human hands, dexterous robot hands, and parallel grippers. Finally, we use domain-adversarial learning to reduce embodiment-specific information in the learned visual representation. Experiments on hand-action prediction and real-world dexterous manipulation show consistent improvements over baselines, effective zero-shot human-to-robot skill transfer, generalization to unseen objects and environments, and data-efficient few-shot adaptation.