ZETA: A Controlled Study of Zero-Shot Cross-Embodiment VLA Transfer for Tabletop Manipulation

📅 2026-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过分析状态-动作表示、预训练多样性、辅助共训练目标和目标实体暴露四个因素,解决了零样本跨实体VLA传输问题,以提高桌面操作的泛化能力。
📝 Abstract
Zero-shot generalization to unseen embodiments is important for generalizable vision-language-action (VLA) models as robot hardware evolves and task-specific data collection remains costly. However, a systematic understanding of this problem remains limited, in part because the literature lacks a unified zero-shot transfer definition and controlled evaluation settings that isolate embodiment changes from differences in tasks, scenes, or protocols. To address this gap, we first distinguish strict zero-shot transfer, where the target embodiment is absent from all training data, from pretrain-exposed zero-shot transfer, where it appears only during pretraining. We then introduce a controlled benchmark spanning 14 held-out target embodiments across simulation and real-world validation. Within this framework, we conduct a controlled analysis of four factors: state-action representations, pretraining embodiment diversity, auxiliary co-training objectives, and target-embodiment exposure. Experimental results show that local end-effector (EEF) state-action representations, the source embodiment diversity, and auxiliary co-training improve cross-embodiment transfer by around 15, 18, and 7 percentage points, respectively. We further find that adding only 5% target-embodiment data during pretraining improves average target-embodiment progress by 13.4 percentage points, showing that strict and pretrain-exposed zero-shot transfer are distinct and should be reported separately. Together, these findings provide practical guidance for evaluating and improving cross-embodiment VLA transfer in stationary tabletop manipulation with two-finger grippers, while motivating future investigation of broader settings including mobile-base control, dexterous hands, and long-horizon tasks.
Problem

Research questions and friction points this paper is trying to address.

Zero-shot Generalization
Cross-embodiment Transfer
Vision-Language-Action Models
Tabletop Manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-shot transfer
Cross-embodiment VLA
State-action representations
Pretraining embodiment diversity
Auxiliary co-training
💼 Related Jobs
No related jobs found.
M
Mi Yan
Galbot; CFCS, School of CS, Peking University
W
Wenhao Zhang
Galbot; Peking University
Z
Zhiqi Zhang
Galbot; Peking University
Yu Peng
Yu Peng
Beihang University
Virtual SurgeryPhysics-based animation
T
Tangxinyu Wang
Galbot; Peking University
L
Lingfei Zhai
Galbot; Peking University
Jiayi Su
Jiayi Su
Northeastern University
HCIHealth
S
Shengliang Deng
Galbot; The University of Hong Kong
L
Lin Peng
Galbot; Beihang University
Y
Yaowei Liu
Galbot; Peking University
Yuxing Chen
Yuxing Chen
Peking University
Embodied AI
Z
Zhiyuan Wei
Galbot; Peking University
Jilong Wang
Jilong Wang
Galbot (Galaxy General Robot Co., Ltd.)
RoboticsReinforcement learningMachine Learning
Jiayi Chen
Jiayi Chen
Peking University
Robotics3D Vision
J
Jiangran Lyu
Galbot; CFCS, School of CS, Peking University
Z
Zhizheng Zhang
Galbot
He Wang
He Wang
Assistant Professor of Computer Science, Peking University
Embodied AIComputer VisionRobotics