EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决长时任务中视觉-语言-动作模型的局限,提出EmbodiedSkills框架,通过统一接口协调感知、规划、执行和验证,实现高效任务执行。
📝 Abstract
Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but long-horizon tasks require more than action prediction. An agent must coordinate perception, planning, execution, progress verification, and recovery as the physical state evolves. An action prediction or a model-generated skill decision does not, by itself, guarantee that the proposed operation is valid in the current state or that its outcome will be verified. We propose EmbodiedSkills, a unified framework that treats each skill decision as an execution proposal: the runtime checks its prerequisites before execution and verifies the outcome afterward. A shared executable-skill interface connects high-level skill selection, bounded low-level VLA execution, and post-action verification within a single agent loop. Because this interface remains fixed, low-level VLA policies can be replaced or adapted without changing the agent loop. The interface also records planning, execution, verification, and recovery events as structured trajectories, which provide supervision for individual components and can support optional online adaptation when interactive feedback is available. We instantiate EmbodiedSkills with Qwen3-VL and OpenPI/pi0.5 on RoboTwin 2.0 and LIBERO. Task-adapted low-level VLA policies achieve an average success rate of 86.20% across 50 RoboTwin 2.0 tasks and 97.40% across the four LIBERO suites. These results establish the execution performance of the task-adapted low-level VLA policies used in EmbodiedSkills. On four memory-dependent RMBench tasks, the same task-adapted execution approach achieves 12.5% average success. The framework provides a trainable and inspectable agent layer for turning these policies into closed-loop embodied systems.
Problem

Research questions and friction points this paper is trying to address.

Vision-language-action
long-horizon tasks
execution verification
skill coordination
closed-loop systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Framework
Execution Proposal
Shared Executable-Skill Interface
Online Adaptation
Structured Trajectories
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Wei Wang
Wei Wang
Zhejiang University
Wireless Communications and Networks
W
Wenqiao Zhang
College of Computer Science and Technology, Zhejiang University
Y
Yutong Lin
College of Computer Science and Technology, Zhejiang University
Yuqian Yuan
Yuqian Yuan
PhD student, Zhejiang University
Computer VisionMachine Learning
Tianwei Lin
Tianwei Lin
Zhejiang University
MLLMs
J
Jinhao Mao
College of Computer Science and Technology, Zhejiang University
Z
Zhenxuan Fan
College of Computer Science and Technology, Zhejiang University
M
Mingjian Gao
College of Computer Science and Technology, Zhejiang University
Yang Dai
Yang Dai
Shenzhen Institutes of Advanced Technology,Chinese Academy of Sciences
perovskitesmemristor
Wentong Li
Wentong Li
Nanjing University of Aeronautics and Astronautics
Computer VisionMachine LearningVision-Language ModelRobotics
Z
Zheqi Lv
Cornell University
Z
Zheng Dong
Universal Ubiquitous AI Co., Ltd.
Y
Yingjie Niu
Hangzhou DEEP Robotics Technology Co., Ltd.
Jiaqi Zhu
Jiaqi Zhu
National University of Singapore
Anomaly DetectionGenerative AIData AnalyticsHealthcare
J
Jun Xiao
College of Computer Science and Technology, Zhejiang University
C
Chao Li
Hangzhou DEEP Robotics Technology Co., Ltd.
Y
Yueting Zhuang
College of Computer Science and Technology, Zhejiang University