Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出ACES框架,通过执行代理工件来评估技能包对企业任务的实际帮助,解决了仅通过结构、风格和安全审查无法评估部署效果的问题。
📝 Abstract
Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and workflow packages must be reviewed with evidence rather than prose. Current gates often scan these artifacts for structure, style, and security, but they do not answer the deployment question: does the capability package help a live agent complete enterprise tasks under the same model, sandbox, and grading policy? We present ACES (Agentic Continuous Evaluation of Skills), a repository-native framework for evaluating skills and product capability packages as executable agent artifacts. ACES runs paired live trials with and without a target skill, normalizes trajectories into the Agent Trajectory Interchange Format (ATIF), grades six default runtime metrics, and reports Skill Lift: the target skill's added value for a fixed task, harness, workspace, and scorer. The same protocol supports product-owned task suites that compare baseline, skill, bundle, team-skill, and plugin targets. On 145 real skills from internal enterprise repositories and public catalogs, scan-only gates surface useful authoring issues but measure complementary facets (structural versus LLM-judge Spearman $ρ= 0.14$). Across 947 scored paired cases from 58 of 64 production skills and four primary harnesses, mean composite Skill Lift is 0.2134 (95\% paired-case CI [0.1967, 0.2301]); mean outcome-only lift, the average of accuracy and goal accuracy, is 0.1799. Composite lift is positive in 72.8\% of paired cases. The largest process-metric gains appear in skill execution, behavior check, and skill efficiency---signals about discovery, routing, workflow following, and tool use that document scans cannot observe. An open-source implementation of the methodology is available in NVIDIA SkillEvaluator.
Problem

Research questions and friction points this paper is trying to address.

Enterprise Agent Programs
Skill Evaluation
Production Deployment
Capability Package
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Continuous Evaluation of Skills
Skill Lift
Agent Trajectory Interchange Format
Runtime Metrics
🔎 Similar Papers
2024-04-08Findings of the Association for Computational Linguistics ACL 2024Citations: 0
C
Christopher Kevin
NVIDIA
N
Narendran Raghavan
NVIDIA
J
Jean-Francois Puget
NVIDIA
R
Roshni Malani
NVIDIA
M
Meghana Puvvadi
NVIDIA
M
Moshe Abramovitch
NVIDIA
M
Mohit Gupta
NVIDIA
Rama Akkiraju
Rama Akkiraju
IBM
AIBusiness process management
S
Subodh Prabhu
NVIDIA
Y
Yogesh Dangi
NVIDIA
W
Wei Luo
NVIDIA
S
Seong Hee Lee
NVIDIA