Grounded Checklist Partial Credit for Agent Skill Trajectories

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决语言模型代理在长周期任务中评估不足的问题,提出了一种基于人类规则和大语言模型的细化评分方法(GCPC),以更准确地评估代理技能轨迹。
📝 Abstract
Language-model agents increasingly tackle long-horizon tasks in interactive environments, yet their evaluation commonly relies on task-level success rates by reducing an entire execution trajectory to whether the task passes an official verifier. This binary score hides partial progress and is particularly limited for procedural agent skill evaluations, since a skill can alter execution without changing the final outcome. While checklists provide finer-grained evaluation by scoring individual task requirements, costly manual authoring and unreliable automatic generation make trustworthy evaluation difficult to scale. To address these challenges, we introduce Grounded Checklist Partial Credit (GCPC), a human-governed and LLM-instantiated partial-credit evaluation of agent trajectories. Humans define reusable rules once, from which an LLM instantiates a task-specific checklist grounded in the task instruction and official verifier. To keep judgment tied to evidence, a judge scores each item from execution log evidence alone and abstains when evidence is missing. A separate scripted step then applies the official verifier outcome to the score. Across a 4,455-trajectory, deduplicated SkillsBench evaluation population, GCPC better discriminates official PASS and FAIL outcomes than holistic judging on the shared subset (AUC 0.689 vs. 0.619). Human evaluation on 96 trajectories from 12 tasks shows that GCPC aligns more closely with human assessments of progress. Applied to 1,946 matched with/without-skill pairs, GCPC exposes the effects hidden by pass@1: among 879 pairs whose binary outcome does not change, 20.9% improve by more than 0.10 while 18.7% regress by the same margin. The GCPC pipeline also transfers to Terminal-Bench and SWE-bench, demonstrating applicability beyond skill-conditioned evaluation.
Problem

Research questions and friction points this paper is trying to address.

language-model agents
long-horizon tasks
task-level success rates
checklists
partial-credit evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Grounded Checklist Partial Credit
Agent Skill Trajectories
Task Evaluation
Partial-Credit Scoring
Human-Governed Rules
💼 Related Jobs
No related jobs found.
S
Suliu Qin
Xi’an Jiaotong-Liverpool University
Lu Yin
Lu Yin
Asst. Professor, CS@University of Surrey & Researcher Fellow, CS@TU/e
AI EfficiencyAI for ScienceRobustnessLLM
X
Xilu Wang
University of Surrey