Behavior-Skill: A Fine-Grained Benchmark for Evaluating Vision-Language-Action Policies in Long-Horizon Tasks

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决长时任务中多技能执行难题,提出Behavior-Skill基准,通过细粒度技能实例评估视觉-语言-动作策略,提供中间状态恢复和成功条件以独立评测。
📝 Abstract
Reliable execution of long-horizon mobile manipulation tasks remains challenging because overall task success depends on the successful completion of multiple constituent skills. Existing benchmarks, however, still rely primarily on full-task rollouts and aggregate task-level metrics, making intermediate failures difficult to observe and analyze. We present Behavior-Skill, a benchmark that reformulates the learning and evaluation of long-horizon tasks around executable constituent skills. It contains 235,492 skill instances from 10,000 demonstrations across 50 household tasks and 34 semantic skill categories. Each instance pairs a skill instruction with an aligned observation-action segment, and is further associated with a restorable intermediate state and a skill success condition to enable independent evaluation under valid preconditions. We further introduce trajectory-level and skill-level metrics to characterize policy capability beyond aggregate task success. Extensive experiments across representative VLA policies including pi0.5 and GR00T on the complete 50-task benchmark show that failures are highly non-uniform across skills, with contact-rich manipulation skills forming persistent bottlenecks. These results demonstrate that Behavior-Skill complements full-task evaluation by exposing intermediate capability profiles for analyzing and improving long-horizon VLA policies. Behavior-Skill is publicly available at https://github.com/nubot-nudt/Behavior-Skill.
Problem

Research questions and friction points this paper is trying to address.

long-horizon tasks
mobile manipulation
constituent skills
benchmark
intermediate failures
Innovation

Methods, ideas, or system contributions that make the work stand out.

Behavior-Skill
long-horizon tasks
skill-level metrics
intermediate states
executable constituent skills
C
Chunyun Ma
College of Intelligence Science and Technology and National Key Laboratory of Equipment State Sensing and Smart Support, National University of Defense Technology, Changsha, China
Lun Luo
Lun Luo
Zhejiang University
SLAMPlace Recognition
X
Xingjian Luo
The Chinese University of Hong Kong, Hong Kong, China; XPeng Inc., Guangzhou, China
X
Xiexing Feng
XPeng Inc., Guangzhou, China
H
Hang Zhang
XPeng Inc., Guangzhou, China
W
Wei Liu
XPeng Inc., Guangzhou, China
Feng Qiao
Feng Qiao
Washington University in St. Louis
Computer VisionArtificial IntelligenceAutonomous Driving
Y
Yaonan Wang
Hunan University, Changsha, China
Huimin Lu
Huimin Lu
National University of Defense Technology
Robot VisionMulti-robot CoordinationRobot SoccerRobot Rescue
Xieyuanli Chen
Xieyuanli Chen
Associate Professor, NUDT, China
RoboticsSLAMLocalizationLiDAR PerceptionRobot Learning