Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

๐Ÿ“… 2026-09-03
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
ๆœฌๆ–‡ๆๅ‡บAdaRoboVLGๆก†ๆžถ๏ผŒ้€š่ฟ‡ๅฏ็ป„ๅˆ็š„ๅŸบ็ก€ๆจกๅž‹ๅ…ˆ้ชŒๅ’Œ้€š็”จๆŠ“ๅ–ๅˆๆˆๆ–นๆณ•่งฃๅ†ณไธๅŒๆœบๆขฐๆ‰‹้—ด็š„้€š็”จๆŠ“ๅ–้—ฎ้ข˜ใ€‚
๐Ÿ“ Abstract
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy that generates and evaluates physically feasible grasp candidates through explicit kinematic mapping and force-closure-based stability estimation, while offloading task-dependent understanding to specialized foundation-model modules. These modules provide composable priors that are integrated into the grasp synthesis process, enabling contextually adaptive grasp synthesis without retraining the underlying grasp policy. Through extensive simulation and real-world experiments, we demonstrate that (i) the base policy exhibits efficient learning and strong cross-hand generalization, (ii) the framework effectively incorporates spatial, cognitive, and temporal priors to address three representative grasping challenges without compromising grasp synthesis performance compared to state-of-the-art methods, and (iii) these priors can operate jointly to enable functional grasping in cluttered and dynamic environments. These results indicate that decoupling physical grasp synthesis from task-dependent understanding provides a scalable paradigm for robotic grasping, allowing future advances in foundation models to be directly translated into improved grasp capabilities without redesigning or retraining the underlying grasp policy. Supplementary videos are available at https://adarobovlg.github.io/
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Grasp
generalizable grasp synthesis
cross-hand generalization
task-adaptive
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Vision-Language Grasping
Composable Foundation Priors
Generalizable Grasp Synthesis
Task-Dependent Understanding
Kinematic Mapping
๐Ÿ”Ž Similar Papers
๐Ÿ’ผ Related Jobs
No related jobs found.
S
Sixu Yan
School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China; also with Hubei Automation Institute, Wuhan 430071, China
S
Shikang Wang
School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China
B
Binhua Huang
School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China
X
Xuanlai Tang
KEENON Robotics Co., Ltd., Shanghai 201206, China
G
Guohua Fan
Suzhou Silicon Era Intelligent Technology Co., Ltd., Suzhou 215131, China
Fan Huang
Fan Huang
Suzhou Zhichuang Xinwei Technology Co., Ltd., Suzhou 215123, China; and College of Materials, Xiamen University, Xiamen 361005, China
Haoxuan Li
Haoxuan Li
University of Electronic Science and Technology of China
MultimediaInformation RetrievalEvent Forecasting
Y
Yongkang Li
School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan 430074, China
Yuhan Li
Yuhan Li
PhD Student in School of Mathematical Sciences, Queen Mary University of London
Network ScieneDigital EpidemiologyComplex Networks
B
Bencheng Liao
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China; and ByteDance, Beijing 100098, China
Zeyu Zhang
Zeyu Zhang
Gaoling School of Artificial Intelligence, Renmin University of China
LLM-based AgentResponsible RecSysCausal Learning
Wenyu Liu
Wenyu Liu
Huazhong University of Science and Technology
Compuetr visionArtificial intelligence
Hangxin Liu
Hangxin Liu
Beijing Institute for General Artificial Intelligence (BIGAI)
RoboticsLocalizationSensors
Xinggang Wang
Xinggang Wang
Professor, Huazhong University of Science and Technology
Artificial IntelligenceComputer VisionAutonomous DrivingObject DetectionObject Segmentation