VTAO-BiManip: Masked Visual-Tactile-Action Pre-training with Object Understanding for Bimanual Dexterous Manipulation

📅 2025-01-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address the challenges of multimodal perception and coordinated control in dexterous bimanual manipulation with dual-arm robots, this paper proposes a masked multimodal pretraining framework that jointly encodes vision, continuous tactile signals (introducing the first hand-motion signal modeling), and action sequences to predict object 6D pose and dimensions, thereby guiding curriculum-based reinforcement learning. Our method features three key innovations: (1) cross-modal masked reconstruction for unified visual-tactile-action representation; (2) a two-stage curriculum RL strategy enabling stable acquisition of multiple sub-skills; and (3) a geometry-aware sim-to-real transfer policy. Evaluated on a bottle-cap twisting task, our approach achieves over 20% higher success rates than state-of-the-art vision-tactile pretraining methods in both simulation and real-world settings, marking the first demonstration of human-hand-level bimanual dexterity in robotic manipulation.

Technology Category

Application Category

📝 Abstract
Bimanual dexterous manipulation remains significant challenges in robotics due to the high DoFs of each hand and their coordination. Existing single-hand manipulation techniques often leverage human demonstrations to guide RL methods but fail to generalize to complex bimanual tasks involving multiple sub-skills. In this paper, we introduce VTAO-BiManip, a novel framework that combines visual-tactile-action pretraining with object understanding to facilitate curriculum RL to enable human-like bimanual manipulation. We improve prior learning by incorporating hand motion data, providing more effective guidance for dual-hand coordination than binary tactile feedback. Our pretraining model predicts future actions as well as object pose and size using masked multimodal inputs, facilitating cross-modal regularization. To address the multi-skill learning challenge, we introduce a two-stage curriculum RL approach to stabilize training. We evaluate our method on a bottle-cap unscrewing task, demonstrating its effectiveness in both simulated and real-world environments. Our approach achieves a success rate that surpasses existing visual-tactile pretraining methods by over 20%.
Problem

Research questions and friction points this paper is trying to address.

Bimanual Manipulation
Robotics
Fine Motor Skills
Innovation

Methods, ideas, or system contributions that make the work stand out.

VTAO-BiManip
Multi-modal Sensory Integration
Bimanual Dexterity Improvement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhengnan Sun
College of Control Science and Engineering, Zhejiang University, Hangzhou, 310027, China
Z
Zhaotai Shi
College of Control Science and Engineering, Zhejiang University, Hangzhou, 310027, China
Jiayin Chen
Jiayin Chen
The University of Hong Kong
human factors engineeringhealth informaticsvirtual reality and rehabilitationserious games
Q
Qingtao Liu
College of Control Science and Engineering, Zhejiang University, Hangzhou, 310027, China
Y
Yu Cui
College of Control Science and Engineering, Zhejiang University, Hangzhou, 310027, China
Q
Qi Ye
College of Control Science and Engineering and the State Key Laboratory of Industrial Control Technology, Zhejiang University, and also with the Key Lab of CS&AUS of Zhejiang Province
J
Jiming Chen
College of Control Science and Engineering, Zhejiang University, Hangzhou, 310027, China