The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents
This study addresses a critical yet overlooked issue in large language model (LLM) agents: while procedural skills improve average task success rates, they often induce “regression”—causing previously solvable tasks to fail. Through controlled experiments on nearly 6,000 office automation tasks, this work quantifies and disentangles the dual effects of skill integration, introducing the concept of a “regression tax.” The findings reveal that skill reliability hinges more on grounding and verification mechanisms than on procedural logic itself; the superiority of optimal skills stems primarily from their lower regression rates; and most regression failures can be mitigated through enhanced verification. These insights establish a new paradigm for skill design, grounded in empirical evidence and emphasizing robustness over mere capability expansion.