CometVLA: Co-Training on an Embodied Data Pyramid towards Physical Understanding

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过构建与机器人动作数据严格对齐的CometData和CometBench,并采用全局动作先验令牌进行协同训练,解决了视觉-语言-动作模型在需要物理常识的操作任务中表现脆弱的问题。
📝 Abstract
Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Current physical VQA data is typically disembodied and misaligned with robot action domains. Egocentric videos are used only as auxiliary pre-training. It remains unclear whether improved VLM physical understanding actually benefits downstream action generation. Therefore, we present CometVLA to close this gap. We construct CometData and CometBench, an embodied physical VQA corpus and benchmark strictly aligned with the robot's action data and embodiment. We introduce Global Action Prior (GAP) tokens, a compact learnable bottleneck that isolates task-agnostic motion regularities and lets the action head consume physical commonsense without corrupting the pre-trained VLM backbone. We co-train CometVLA across the embodied data pyramid, spanning teleoperation, simulation, egocentric trajectories, and VQA layers. On real-world manipulation tasks and RoboTwin simulation, CometVLA consistently outperforms strong VLA baselines. Correlation analysis shows that stronger VLM performance on CometBench indicates higher VLA success rates. Results demonstrate that physical understanding pre-training genuinely benefits downstream manipulation.
Problem

Research questions and friction points this paper is trying to address.

Vision-language-action
physical commonsense
embodied data
action generation
manipulation tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Co-Training
Embodied Data Pyramid
Global Action Prior (GAP)
Physical Commonsense
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hanwen Wan
JD Explore Academy, China; School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China; Shenzhen Institute of Artificial Intelligence and Robotics for Society, China
D
Dafeng Chi
JD Explore Academy, China
L
Linbo Zhai
JD Explore Academy, China; South China University of Technology, China
T
Tianao Shen
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China; Shenzhen Institute of Artificial Intelligence and Robotics for Society, China
Yuzheng Zhuang
Yuzheng Zhuang
Senior Researcher @ Huawei Noah's Ark Lab
Reinforcement LearningOptimizationAutonomous DrivingCommunication
T
Tianle Zhang
JD Explore Academy, China
Peidong Liu
Peidong Liu
Westlake University
3D computer visionRobotics
Liang Lin
Liang Lin
Fellow of IEEE/IAPR, Professor of Computer Science, Sun Yat-sen University
Embodied AICausal Inference and LearningMultimodal Data Analysis
X
Xiaoqiang Ji
School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China; Shenzhen Institute of Artificial Intelligence and Robotics for Society, China; School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China