Gripper-aware Vision Language Action Models

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决现有视觉语言动作模型忽视夹爪类型差异的问题,通过构建多夹爪数据集MiGA和提出GVLA方法来提升机器人对不同夹爪的适应性和操作性能。
📝 Abstract
Vision language action models (VLAs) have advanced general purpose robotic grasping and manipulation by enabling robots to interpret visual observations and natural language instructions to generate executable action sequences. However, existing VLAs often implicitly assume gripper invariance, despite grasping strategies being inherently embodiment-dependent. Different gripper types, such as parallel-jaw and suction, usually require distinct interaction strategies to achieve the same grasping objective. Moreover, current datasets for VLAs predominantly rely on parallel-jaw grippers, limiting gripper-aware learning. To address this gap, we introduce MiGA, a multi-gripper-aware dataset spanning five distinct gripper types across multiple robots with 103,000 demonstrations, explicitly capturing strategy divergence under shared task objectives. We further propose GVLA, which combines a new multi-gripper tokenizer with adapter-based policy routing. Our new gripper encoding induces structured embedding information that balances parameter sharing and strategy differentiation, while layer-wise probing confirms meaningful gripper-conditioned representations for VLAs. Intensive experiments in both simulation and real-world robots show that our GVLA outperforms the current baselines across evaluated settings. Our method also improves zero-shot generalization or few-shot adaptation to new objects or unseen tasks, and enable more efficient gripper adaptation.
Problem

Research questions and friction points this paper is trying to address.

Vision Language Action Models
Gripper Invariance
Embodiment-Dependent Strategies
Multi-Gripper Dataset
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-gripper-aware dataset
adapter-based policy routing
gripper encoding
💼 Related Jobs
No related jobs found.
Hanyi Zhang
Hanyi Zhang
University of Heidelberg
Deep LearningBiomedical Imaging
Z
Zihong Luo
University of Liverpool, UK
T
Tianyu Li
University of Liverpool, UK
Khang Nguyen
Khang Nguyen
UIT
Artificial IntelligenceComputer VisionTimetabling
B
Basu Hela
Indian Institute of Science, India
S
Shreyas Kumar
Indian Institute of Science, India
N
Ngoc Duy Tran
Indian Institute of Science, India
Feng Dai
Feng Dai
Institute of Computing Technology, Chinese Academy of Sciences
video coding and processingcomputational imaging
Charith Munasinghe
Charith Munasinghe
Researcher at Zurich University of Applied Science
Cloud RoboticsArtificial IntelligenceMicro Aerial VehiclesRobotics
Jorge Peña Queralta
Jorge Peña Queralta
Research Group Leader @ ZHAW, Adjunct Professor at University of Turku
roboticssocial navigationmulti-robot systemsassistive roboticsshared control
Giovanni Toffetti
Giovanni Toffetti
Zürcher Hochschule für Angewandte Wissenschaften, Switzerland
Khoa Vo
Khoa Vo
Postdoctoral Fellow at the University of Arkansas, USA
Vision Language ModelComputer VisionDeep Learning
Ngan Le
Ngan Le
University of Arkansas
Artificial IntelligenceMachine LearningComputer Vision
Ravi Prakash
Ravi Prakash
Assistant Professor, Indian Institute of Science
Robot LearningRoboticsOptimization and Optimal Control
Quan Vuong
Quan Vuong
Physical Intelligence
Reinforcement LearningComputer Vision
Tung D. Ta
Tung D. Ta
The University of Tokyo
RoboticsHuman Computer InteractionDigital Fabrication
Long Hu
Long Hu
Associate Professor of Computer Science, Huazhong University of Science and Technology
Edge ComputingBig DataAffective ComputingDeep Reinforcement Learning
Anh Nguyen
Anh Nguyen
University of Liverpool
Robotic VisionMachine LearningRobotics
Baoru Huang
Baoru Huang
University of Liverpool; Imperial College London
RoboticsComputer visionSurgical visionImage-Guided Intervention