Scaling Automatic Research Agents via World Models

πŸ“… 2026-08-12
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the high cost of real-environment execution that hinders autonomous scientific agents in reinforcement learning. To overcome this limitation, we introduce world models into this setting for the first time and propose a World Model Reinforcement Learning (WMRL) framework. WMRL integrates online debiasing and inverse-variance denoising mechanisms to simultaneously correct reward bias and suppress noise, with theoretical guarantees. This approach substantially alleviates the environment interaction bottleneck, accelerating training by 3–4Γ—. The resulting 4B/9B-parameter agents outperform significantly larger open-source models (48B/120B) on established benchmarks while preserving performance. Furthermore, the trained agents successfully transfer to embodied vision-language-action (VLA) policy fine-tuning, demonstrating the framework’s generality and effectiveness.
πŸ“ Abstract
Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring this goal within reach, as modern LLMs show the capability to independently implement solutions and learn from the execution outcomes. Behind these gains, post-training (especially RL) plays a central role. In this paper, we identify a fundamental tension when scaling RL for these agents: the two components of every AutoResearch trajectory (agent generation and environment execution) scale in very different manners, since all generation shares compute through batching, while each execution occupies its exclusive sandbox and real machine time. As a result, the environment execution dominates the training cost and becomes the bottleneck as trajectories grow. To resolve this tension, we propose World Model RL (WMRL), which replaces environment execution with a world model to remove this bottleneck. Additionally, the world model can be imperfect, as its rewards are corrupted by bias and noise. Therefore, we further equip WMRL with two mitigations, Online Debiasing and Inverse-Variance Denoising, which offset the bias and suppress the noise respectively. Theoretically, we prove that both mitigations of WMRL strictly improve the convergence guarantee. Empirically, WMRL accelerates training by 3-4x on various tasks at different agent scales, while exceeding the performance of standard RL baselines. Moreover, our post-trained 4B and 9B agents outperform much larger open-weight agents of 48B and 120B on held-out benchmarks. Beyond AutoResearch, WMRL also transfers to post-training embodied VLA policies, which demonstrates the generalizability of our method.
Problem

Research questions and friction points this paper is trying to address.

AutoResearch
reinforcement learning
training bottleneck
environment execution
scaling
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Model RL
AutoResearch
Online Debiasing
Inverse-Variance Denoising
Scalable Reinforcement Learning