ARISE-RL: Agentic Rubric-Grounded Iterative Self-Evolution with Reinforcement Learning

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出ARISE-RL框架,通过生成器和求解器的协同进化解决开放性任务中强化学习面临的奖励不稳定问题,并引入RG-SED方法优化策略学习。
📝 Abstract
Training open-ended agents via reinforcement learning (RL) is hindered by the lack of verifiable gold answers and scalable rubrics. Moreover, even near the model's capability boundary, long-horizon open-ended agentic tasks often yield brittle and unstable rewards, resulting in weak or noisy rollout contrast that obscures fine-grained optimization signals for group-based policy learning. To address these challenges, we propose ARISE-RL, a novel full-cycle self-evolution framework that couples a task/rubric Generator and a reasoning Solver through rubric-mediated co-evolution. The Generator grounds tool-related rubric criteria in real tool observations and is rewarded for producing valid, intermediate-difficulty tasks aligned with the Solver's evolving capability boundary. The Solver, in turn, learns from fine-grained rubric satisfaction signals through multi-step reasoning and tool use. We further introduce Reward-Gated Self-Evolution Distillation (RG-SED), which selectively distills a memory-augmented variant of the same policy back into itself only when the memory yields empirical reward improvement, thereby reducing distribution mismatch and avoiding blind imitation of noisy guidance. Finally, to support rigorous evaluation, we present ECR-Bench, an expert-calibrated rubric benchmark suite covering single-tool deep research and multi-tool travel planning. Extensive experiments demonstrate that ARISE-RL consistently achieves robust and stable overall state-of-the-art performance across all evaluated benchmarks.
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
open-ended agents
verifiable gold answers
scalable rubrics
unstable rewards
Innovation

Methods, ideas, or system contributions that make the work stand out.

ARISE-RL
Rubric-Mediated Co-Evolution
Reward-Gated Self-Evolution Distillation
RG-SED
🔎 Similar Papers
F
Fanrui Zhang
Alibaba ATH Token Foundry
R
Ruixue Ding
Alibaba ATH Token Foundry
Q
Qiang Zhang
Alibaba ATH Token Foundry
X
Xi Chen
Alibaba ATH Token Foundry
Boli Chen
Boli Chen
University College London
Systems and ControlOptimizationSmart Cities
Shihang Wang
Shihang Wang
DAMO Academy, Alibaba Inc.
Natural Language Processing
H
Hongmin Zhan
J
Jinxin Bian
Hema, Alibaba Group
X
Xingchao Li
Hema, Alibaba Group
P
Peijin Zheng
Hema, Alibaba Group
H
Hao Cheng
Hema, Alibaba Group
Pengjun Xie
Pengjun Xie
Alibaba Group
NLP/IR/ML
Kaipeng Zhang
Kaipeng Zhang
Shanghai AI Laboratory
LLMMultimodal LLMsAIGC
J
Jiawei Liu
Z
Zheng-Jun Zha