SETA: Scaling Environments for Terminal Agents
This work addresses the challenge of limited large-scale, diverse, and verifiable training environments for terminal-based reinforcement learning (RL). It introduces SETA, a framework featuring two pipelines—SETA-Synth and SETA-Evol—that enable the first scalable and verifiable automatic generation of terminal RL environments. The framework incorporates a unified verification mechanism and a difficulty-adaptive evolution strategy, yielding SETA-Env, an open dataset comprising over 4,500 tasks. By integrating instruction synthesis, environment construction, and automated validation, the authors train agents using the GRPO algorithm on Qwen3-8B and DeepSeek-V4-Flash models. On Terminal-Bench 2.0, these agents achieve 12% pass@1 (state-of-the-art for 8B models) and 43% pass@1 (a 3% improvement), with pass@5 reaching 58% (a 4% gain).