Skill Following: Evaluating Actual Skill Use in Retrieval-Enabled LLM Agents

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入RAE方法,解决了评估LLM在实际使用外部技能时的真实效果问题,揭示了聚合指标可能造成的误导。
📝 Abstract
Large Language Model (LLM) agents increasingly rely on external skills, yet standard evaluations obscure whether retrieving these skills actually helps. Aggregate metrics often compare retrieved versus non-retrieved tasks, introducing severe selection bias and failing to isolate the true effect of skill use. To measure this actual-use capability-which we formalize as Skill Following (SF)-we introduce the Retrieval-Invoked Actual-Use Effect (RAE). RAE computes the same-task outcome difference between matched skill-enabled and skill-disabled executions, conditioned exclusively on tasks where the agent actively retrieved a skill. Evaluating 17 LLMs across coding and mathematical domains, we uncover a stark evaluation paradox: models frequently show positive aggregate retrieval lift but negative RAE. On MBPP+, multiple models that appear to benefit system-wide actually harm their own performance on the exact tasks where retrieval occurred. These findings demonstrate that aggregate averages can create a misleading illusion of tool-use proficiency, whereas RAE directly measures whether the retrieval-to-answer pipeline genuinely rescues more outcomes than it harms.
Problem

Research questions and friction points this paper is trying to address.

Skill Following
Retrieval-Enabled LLM Agents
Evaluation Paradox
Retrieval-Invoked Actual-Use Effect
Selection Bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

Skill Following
Retrieval-Invoked Actual-Use Effect
external skills
evaluation bias
LLM agents
🔎 Similar Papers
💼 Related Jobs
No related jobs found.