DynSTEER: Dynamic Stage-wise Trajectory Evaluation and Execution-time Review for Agents
为解决长任务执行中评估方法的局限性,DynSTEER通过动态阶段式轨迹评估和容错里程碑图来提高评估准确性并减少资源浪费。
为解决长任务执行中评估方法的局限性,DynSTEER通过动态阶段式轨迹评估和容错里程碑图来提高评估准确性并减少资源浪费。
This study addresses the challenge of fixed environment generation strategies failing to adaptively align with learning frontiers in reinforcement learning for terminal agents. We propose Envs-FORGE, a novel framework introducing a seed-wise dynamic environment synthesis mechanism guided by learning frontiers. This approach translates validator rewards into synthetic actions and optimizes generation policies via mixed-integer linear programming while synchronously reconstructing test environments, retaining only gold-standard data for training. Experiments on Qwen3.5-35B demonstrate that Envs-FORGE improves tb-core Pass@1 by 9.2% and achieves 77.1% on SWE-bench Verified, significantly outperforming fixed-strategy baselines. These results confirm the framework’s effectiveness in overcoming the limitations of conventional prompting strategies by enabling adaptive curriculum generation tailored to evolving model capabilities.
This work addresses performance bottlenecks arising from inefficient communication and computation scheduling when training large language models under constrained hardware resources. The authors propose a unified mixed-integer programming framework that jointly optimizes checkpoint selection, activation placement, recomputation strategies, and overlapping of CPU–GPU–NVMe communications, formulating multidimensional resource scheduling as a mixed-integer problem for the first time. Additionally, they introduce a Hybrid 8-bit operator that integrates 8-bit optimizer state compression with fast gradient clipping to reduce memory usage while preserving computational efficiency. Experiments on training the Qwen3.6-27B model on H800 GPUs demonstrate a peak throughput of 1,361 tokens/s (a 1.24× improvement), requiring only 68.84 GB of GPU memory, achieving 95.42% accuracy, and delivering an effective compute performance of 219.95 TFLOPS.
为解决长任务执行中评估方法的局限性,DynSTEER通过动态阶段式轨迹评估和容错里程碑图来提高评估准确性并减少资源浪费。
This study addresses the challenge of fixed environment generation strategies failing to adaptively align with learning frontiers in reinforcement learning for terminal agents. We propose Envs-FORGE, a novel framework introducing a seed-wise dynamic environment synthesis mechanism guided by learning frontiers. This approach translates validator rewards into synthetic actions and optimizes generation policies via mixed-integer linear programming while synchronously reconstructing test environments, retaining only gold-standard data for training. Experiments on Qwen3.5-35B demonstrate that Envs-FORGE improves tb-core Pass@1 by 9.2% and achieves 77.1% on SWE-bench Verified, significantly outperforming fixed-strategy baselines. These results confirm the framework’s effectiveness in overcoming the limitations of conventional prompting strategies by enabling adaptive curriculum generation tailored to evolving model capabilities.
This work addresses performance bottlenecks arising from inefficient communication and computation scheduling when training large language models under constrained hardware resources. The authors propose a unified mixed-integer programming framework that jointly optimizes checkpoint selection, activation placement, recomputation strategies, and overlapping of CPU–GPU–NVMe communications, formulating multidimensional resource scheduling as a mixed-integer problem for the first time. Additionally, they introduce a Hybrid 8-bit operator that integrates 8-bit optimizer state compression with fast gradient clipping to reduce memory usage while preserving computational efficiency. Experiments on training the Qwen3.6-27B model on H800 GPUs demonstrate a peak throughput of 1,361 tokens/s (a 1.24× improvement), requiring only 68.84 GB of GPU memory, achieving 95.42% accuracy, and delivering an effective compute performance of 219.95 TFLOPS.