Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过引入基于共享文件系统的管理-工作架构,无需额外训练或调整,提高了多模型系统在编码任务中的表现,尤其对某些模型效果显著。
📝 Abstract
Multi-agent large language model systems are widely reported to beat single-model baselines, but the evidence is mixed, and comparisons are usually confounded: pipelines change token budgets, tool calls, and prompts simultaneously, so an aggregate gain rarely reveals what actually helped. We investigate the effect of introducing the manager-worker scaffold over a shared filesystem workspace, with no training and no per-benchmark tuning, measured against the same model answering in a single pass. Across nine models -- five open-weight, spanning 9B to ~2.8T parameters, and four frontier closed models -- on the 100 latest hard LiveCodeBench problems, the scaffold's benefit is real but conditional: large and statistically significant for some (Qwen3.8-27B +23.4, GPT-5.6-Luna +10.6 and GPT-5.6-Terra +8.0, each over five paired passes; Kimi-K3 +30.4 and Minimax-M3 +11.0 over five paired passes with reasoning off, both at $p < 10^{-4}$, and +42 and +12 in a single pass at a 128k cap) and null or negative for others (Qwen3.6-35B -1 to -9 with reasoning off). With the manager, Opus-5 achieves the highest score in the study at 91% in one pass. Running a manager roughly triples the token bill, but it buys accuracy more cheaply than moving to a larger model does: GPT-5.6-Terra with a manager nearly matches Fable 5's single-call accuracy (85.0 against 87.4, $p = 0.59$) at a fifth of the price (\$11.71 against \$61.11 per 100-problem pass, $p < 10^{-4}$), and the Qwen-27B arm does it for \$51.75 on weights anyone can self-host. Our transcript analysis finds several mechanisms behind the gains, of which two recur: context management, in which short worker calls and shared notes organize state and reduce truncation, and problem decomposition. Improvements are modest for large models with reasoning enabled, but larger for some models with reasoning disabled and for smaller models with reasoning enabled.
Problem

Research questions and friction points this paper is trying to address.

multi-agent large language model systems
manager-worker scaffold
coding performance
shared filesystem workspace
context management
Innovation

Methods, ideas, or system contributions that make the work stand out.

manager-worker scaffold
shared filesystem workspace
context management
problem decomposition
💼 Related Jobs
No related jobs found.
V
Victor Gao
Persis Capital Inc.
V
Vida Khosrowshahi
Persis Capital Inc.
A
Ali Khosrowshahi
Persis Capital Inc.
X
Xihao Sun
Persis Capital Inc.
Juhyun Lee
Juhyun Lee
University of Texas at Arlington
Cardiac DevelopmentBiomechanicsOptical Imaging
S
Simon (Sang Won) Lee
Persis Capital Inc.