RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls
This work investigates how coordination content—such as shared states or tool outputs—in multi-agent or memory-augmented large language models consumes tokens within a fixed context budget, thereby displacing task instructions and evidence and degrading performance. To quantify this “task budget displacement” effect, the authors introduce the Roundtable Context Window Test (RCWT), which systematically controls total token budget, content ordering, task type, and evaluation metrics to disentangle the impact of coordination content volume from competition for contextual resources. Through controlled prompting, context window scaling, ablation studies, and experiments across multiple models—including GPT-4.1-mini, Claude Haiku 4.5, and Gemini 2.5 Flash—they find that performance sharply declines when fewer than a few hundred tokens remain for task evidence. However, if the evidence is fully preserved, models can achieve perfect task completion even when coordination content occupies up to 95% of the context window.