EnvCraft: Synthesizing Executable Environments in Agentic RL for Claw-like Agent

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
The paradigm of LLMs has rapidly shifted from passive language interfaces to autonomous Claw-like agents that execute long-horizon tasks across stateful workspaces. While Agentic Reinforcement Learning (Agentic RL) provides a promising path to optimize these agents, its scaling is heavily bottlenecked by the severe scarcity of interactive training environments. Existing synthetic environments are strictly limited to tool-calling endpoints, rendering them insufficient for accommodating the end-to-end real-world demands of claw-like agents. To bridge this gap, we introduce EnvCraft, an automated framework for synthesizing executable environments and scalable training data. Specifically, EnvCraft employs an environment synthesis engine to build sandbox-isolated workspaces, alongside a topology-aware data generation engine to produce coherent task trajectories. Overall, we synthesize 139 interactive environments comprising approximately 20K complex tasks for Agentic RL training. Experiments on Qwen3/3.5 models (8B-32B) show that our method yields gains of up to +11.9% on Claw-style benchmarks and +8.0% on general tool-use benchmarks, with concurrent reductions in inference token cost. The results confirm that synthesized executable environments provide robust and generalizable learning signals for training.
Problem

Research questions and friction points this paper is trying to address.

Agentic RL
Claw-like agents
Interactive training environments
Synthetic environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

EnvCraft
environment synthesis engine
topology-aware data generation
Agentic RL
executable environments
Y
Yirong Zeng
Harbin Institute of Technology, SCIR Lab
S
Shen You
Harbin Institute of Technology, SCIR Lab
J
Jinhang Feng
Peking University
Y
Yufei Liu
Peking University
Xiao Ding
Xiao Ding
Harbin Institute of Technology
Natural Language ProcessingArtificial Intelligence
Yutai Hou
Yutai Hou
Huawei
LLMNLPDialogueAlignmentMeta Learning
H
Hao Cong
Tsinghua University
Y
Yuxian Wang
Huawei Technologies Co., Ltd
W
Wu Ning
Huawei Technologies Co., Ltd
Wang Xu
Wang Xu
Harbin Institute of Technology
natural language processingartificial intelligence
Bibo Cai
Bibo Cai
Harbin Institute Technology
NLP