Self Improvement via Fast Tree-search

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种基于快速树搜索的自我改进框架SIFT,通过减少评估成本和利用LLM作为裁判来提高编码性能,显著降低了资源需求。
📝 Abstract
Coding agents can recursively modify their own implementations, forming a loop of self-improvement. While prior work shows this can boost performance on coding benchmarks, existing approaches are costly and compute-intensive. We introduce a simple, sample-efficient self-improvement framework that significantly improves coding performance under strict budget constraints. We identify evaluation of candidate self-modifications as the main runtime bottleneck since prior approaches estimate their effectiveness by re-running a subset of benchmark tasks with the modified agent, which is time-consuming. We introduce Recursive Self Improvement via Fast Tree-search (SIFT), which augments these downstream task evaluations with an LLM-as-a-judge signal that performs pairwise comparisons between candidate patches, where the win-loss record is aggregated with a regularized Bradley-Terry model, and the resulting strength scores drive rank-based parent sampling inside a lightweight disaggregated tree search. Expensive downstream task evaluations are reserved only for the most promising nodes. Using a fully disaggregated tree search pipeline, the judge scores provide intermediate signal to guide exploration on promising candidate patches without being bottlenecked by slow evaluation runs. SIFT outperforms existing tree-search based self-evolution frameworks on the full Polyglot benchmark with significantly lower resource requirements in terms of CPU hours, wall clock time, and API cost.
Problem

Research questions and friction points this paper is trying to address.

self-improvement
coding agents
budget constraints
evaluation bottleneck
resource requirements
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Self Improvement
Fast Tree-search
LLM-as-a-judge
Bradley-Terry model
disaggregated tree search
🔎 Similar Papers
No similar papers found.