AdvDex: Learning Dexterous Manipulation from Human Demonstrations via Joint-Aligned Actions and Adversarial Learning

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of data scarcity in dexterous manipulation and cross-embodiment generalization by proposing a unified learning framework. By constructing the OmniShare dataset and introducing the JAAS universal action representation, combined with SE(3) pose encoding and domain-adversarial learning, this approach effectively decouples and integrates human-robot demonstration data. Leveraging a Vision-Language-Action model, the proposed method achieves zero-shot human-to-robot skill transfer. Consequently, it significantly enhances cross-embodiment generalization capabilities and few-shot adaptation efficiency, establishing a novel paradigm for universal manipulation in embodied intelligence.
📝 Abstract
Dexterous manipulation is a fundamental capability for embodied intelligence, but scaling it remains difficult because robot demonstrations are expensive to collect and action spaces vary across embodiments. Policies trained on heterogeneous data can also entangle task-relevant visual cues with embodiment-specific appearance, limiting cross-embodiment generalization. We present AdvDex, a unified Vision-Language-Action framework for learning dexterous manipulation from human and robot demonstrations. First, we introduce OmniShare, a large-scale multimodal dataset of human manipulation demonstrations that provides high-quality kinematic supervision and tactile measurements while reducing reliance on robot teleoperation. Second, we propose the Joint-Aligned Action Space (JAAS), a canonical action representation comprising an $\mathrm{SE}(3)$ wrist pose and 15 finger joints, thereby functionally aligning human hands, dexterous robot hands, and parallel grippers. Finally, we use domain-adversarial learning to reduce embodiment-specific information in the learned visual representation. Experiments on hand-action prediction and real-world dexterous manipulation show consistent improvements over baselines, effective zero-shot human-to-robot skill transfer, generalization to unseen objects and environments, and data-efficient few-shot adaptation.
Problem

Research questions and friction points this paper is trying to address.

Dexterous Manipulation
Cross-embodiment Generalization
Heterogeneous Data
Human-to-Robot Transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Joint-Aligned Action Space
Domain-Adversarial Learning
OmniShare Dataset
Cross-Embodiment Generalization
Vision-Language-Action Framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhiyue Zhao
Zhejiang University
J
Jingyi Wu
Fudan University
H
Hairuo Liu
Shanghai Jiao Tong University
Mingyu Liu
Mingyu Liu
Technical University of Munich
Computer VisionDeep Learning
L
Liyang Li
Zhejiang University
H
Hengdi Zhang
Paxini Tech
T
Tong He
Shanghai Innovation Institute
Zhengxue Cheng
Zhengxue Cheng
Assistant Researcher, Shanghai Jiao Tong University
Video and Image CodingComputer VisionImage Quality Assessment