OpenDexGrasp: Open-vocabulary Task-Oriented Dexterous Grasping

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出OpenDexGrasp框架,通过结合视觉-语言上下文和灵巧动作生成,解决开放词汇任务导向的灵巧抓取问题。
📝 Abstract
Dexterous grasp synthesis has advanced rapidly in generating stable and physically plausible hand poses, but real-world manipulation requires grasps that preserve the function implied by the task. We study open-vocabulary task-oriented dexterous grasp generation, where a robot must infer functional intent from free-form language, ground it in multi-view visual observations and object geometry, and generate an executable high-degree-of-freedom grasp. We present OpenDexGrasp, a unified data and generative modeling framework for this setting. OpenDexVerse provides dual-source supervision organized by the Coverage-to-Alignment (C2A) Recipe: OpenDex-Scale offers large-scale semantic and geometric coverage through automatic grasp synthesis and vision-language annotation, while OpenDex-Align supplies high-quality embodied alignment through human teleoperation and category-level transfer. OpenDexGrasp learns a shared perception-action latent representation that couples open-vocabulary vision-language context with dexterous action generation. Affordance grounding and grasp generation provide complementary supervision over this latent space, enabling direct generation of task-consistent dexterous grasps without a separate affordance-to-pose inference stage. Extensive simulation and real-robot experiments demonstrate improved functional alignment, physical feasibility, generalization to unseen categories, and real-world execution success. Additional details and videos are available at https://opendexgrasp.github.io/.
Problem

Research questions and friction points this paper is trying to address.

task-oriented dexterous grasping
open-vocabulary
functional intent
multi-view visual observations
high-degree-of-fedom grasp
Innovation

Methods, ideas, or system contributions that make the work stand out.

open-vocabulary task-oriented dexterous grasping
shared perception-action latent representation
Coverage-to-Alignment (C2A) Recipe
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Jiyao Zhang
Jiyao Zhang
Peking University
Embodied AIRobotics3D Vision
J
Junhan Wang
CFCS, School of CS, PKU, China; PrimeBot; State Key Laboratory of General Artificial Intelligence, BIGAI
T
Tianyu Wang
CFCS, School of CS, PKU, China; State Key Laboratory of General Artificial Intelligence, BIGAI
Z
Zeyuan Chen
CFCS, School of CS, PKU, China; State Key Laboratory of General Artificial Intelligence, BIGAI
A
Anthony Bolten
CFCS, School of CS, PKU, China; PrimeBot
Y
Yitong Peng
CFCS, School of CS, PKU, China
Hao Dong
Hao Dong
Peking University. Associate Professor at Center for Social Research, Guanghua School of Management
KinshipFamilySocial DemographyHistorical DemographySocial Stratification