Experience Funnel: A State-Policy Alternating Loop for Self-Evolving Agents

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究通过结合快速状态适应和慢速策略整合的循环方法,解决了如何将大量任务特定交互经验转化为可复用模型能力的问题。
📝 Abstract
Autonomous agents powered by large language models (LLMs) continuously accumulate experience through interaction, creating an opportunity to improve future behavior through self-evolution. A fundamental challenge is how to transform abundant, task-specific interaction experience into reusable model competence without sacrificing the ability to adapt rapidly to newly observed evidence. Explicit textual states, such as skills and agent harnesses, provide fast, human-readable and editable adaptation, but incur persistent dependence on external context; parametric policies provide compact and reusable competence, but are substantially slower to update. We present \textit{Experience Funnel}, a self-evolving framework that couples fast state adaptation with slow policy consolidation in an alternating loop. Interaction trajectories are first distilled into an explicit textual state, where newly acquired experience can be rapidly incorporated and validated. The framework then selectively identifies state-enabled behavior that remains useful across state revisions and consolidates it into the policy through transition-aware distillation. The updated state--policy pair subsequently generates new rollouts, providing fresh evidence for the next round of state adaptation and policy consolidation. Experiments across diverse agent benchmarks show that \textit{Experience Funnel} consistently improves agent capability over state-only evolution and policy-internalization approaches, while progressively converting useful explicit experience into autonomous policy competence.
Problem

Research questions and friction points this paper is trying to address.

autonomous agents
self-evolution
large language models
experience transformation
adaptation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Experience Funnel
state-policy alternating loop
transition-aware distillation
self-evolving agents
🔎 Similar Papers
No similar papers found.
W
Wenbo Gao
The Hong Kong Polytechnic University, Huawei
Z
Zhaomou Song
Huawei
Z
Zhiyuan Ji
Huawei, Renmin University of China
R
Renxi Liu
Huawei
Xing Li
Xing Li
Huawei Noah's Ark Lab
LLM InferenceTest Time ScalingAgentic AILogic Synthesis
Xianzhi Yu
Xianzhi Yu
Unknown affiliation
AIHPC
Xiaoguang Li
Xiaoguang Li
Noah's Ark Lab,HUAWEI
Question AnsweringInformation RetrievalDialogue Systems
J
James Chung-wai Cheung
The Hong Kong Polytechnic University
Weizhe Lin
Weizhe Lin
University of Cambridge
Natural Language ProcessingAffectie ComputingComputer Vision
Y
Yaoyuan Wang
Huawei