PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出PILOT架构,通过实时指导和自我进化机制解决长周期代理执行过程中经验利用延迟的问题,提高任务执行效率。
📝 Abstract
Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends, so they cannot redirect the active run or immediately apply and validate lessons learned from it. We argue that self-improvement should instead be live, using emerging experience both to redirect the active run and to update the persistent harness. Existing agent architectures do not fully support this goal. Single-agent self-correction combines task execution and trajectory assessment within one context, while subagent delegation separates execution but typically cannot redirect an active subagent. We present PILOT, a supervisor-worker harness for live self-improvement through two coupled mechanisms: (1) live steering lets a separate supervisor redirect or abort the active worker during execution; and (2) live self-evolution distils procedures and failure modes revealed during execution into reusable skills and memory. Across two frozen backbones and three benchmarks, PILOT ranks first in five of six configurations. On Terminal-Bench 2.0, PILOT outperforms counterpart harnesses by up to 9.8 percentage points. In the self-improvement setting, PILOT gains 14.6 points with GLM-5.1 and 12.4 points with Kimi-K2.6. Mean output tokens fall by 42.9% and 47.4%, while successful evaluations per million output tokens rise by 110.3% and 134.0%, respectively.
Problem

Research questions and friction points this paper is trying to address.

Long-horizon Agents
Live Self-Improvement
Experience Utilization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Live Self-Improvement
Supervisor-Worker Harness
Live Steering
Live Self-Evolution
Long-Horizon Agents
Y
Yang Xiao
Y
Yusong Sun
Haoyi Wu
Haoyi Wu
ShanghaiTech University
W
Wenyang Hui
W
Wen Da
Z
Zhaokai Luo
M
Mu Chuan
Yao Hu
Yao Hu
浙江大学
Machine Learning
W
Wenjie Li
Chengyue Jiang
Chengyue Jiang