Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出一种自进化代理框架,通过识别和修正不稳定步骤来解决大语言模型在执行相同任务时的一致性差距问题。
📝 Abstract
Large language model (LLM)-powered agents can be accurate on average yet unreliable in production, a discrepancy that has been observed but remains largely unaddressed. When given the same task five times, a ReAct agent on the AppWorld benchmark using GPT-4.1 succeeds in all five runs only 53% of the time, even though its per-run pass rate averages 77%. We call this 24-point shortfall the consistency gap, and we argue that addressing it is a precondition for trustworthy AI agent deployment. We present a self-evolving agent framework that reduces this gap by identifying unstable, low-consistency steps in agent trajectories and converting them into episodic memory the agent can draw on in future runs. At its core is a Consistency Analyzer that pinpoints where and why a trajectory is likely to flip across executions, and a Guideline Generator that converts the diagnosis into targeted guidelines, committed to memory and injected into future agent executions on similar tasks. On AppWorld with ReAct/GPT-4.1, our framework raises the fraction of tasks that succeed in all five runs by +16 points on same-task evaluation and +13 points on similar-task generalization.
Problem

Research questions and friction points this paper is trying to address.

consistency gap
large language model
unreliable in production
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-evolving agent framework
Consistency Analyzer
Guideline Generator
💼 Related Jobs
No related jobs found.
Evelyn Duesterwald
Evelyn Duesterwald
IBM Research
B
Benjamin Elder
IBM Software Innovation Lab
L
Lilian Ngweta
IBM Software Innovation Lab
Shashanka Ubaru
Shashanka Ubaru
IBM Research
Numerical Linear AlgebraMachine LearningQuantum Algorithms
M
Malgorzata Zimon
IBM Research