EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出EvoRS框架,通过自进化奖励系统解决开放式强化学习中的奖励失效问题,提高任务表现并减少奖励黑客攻击。
📝 Abstract
Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable answers. Yet the policy and reward system form a dynamic feedback loop: as the policy optimizes the current reward, an initially useful reward system may become unreliable due to reward hacking or reduced response discriminability. The reward system should therefore evolve rather than remain fixed during training. Existing dynamic-rubric methods adapt evaluation criteria, but reward failures can also arise from scoring mechanisms or signal composition. We introduce EvoRS, a self-evolving RL framework that evolves the reward system from on-policy experience, representing it as an executable Reward-DAG. Specifically, an agentic designer updates this system from on-policy rollouts and reward traces to maintain train-time reliability. Across writing and roleplay, EvoRS achieves the best quality under all three judges, outperforming the policy by \(2.107\) and \(4.767\) points, respectively, while reducing reward hacking and coverage failures and preserving reward informativeness. Ablations confirm that a comprehensive fixed reward system cannot remain reliable in open-ended tasks and must evolve throughout training.
Problem

Research questions and friction points this paper is trying to address.

open-ended reinforcement learning
reward system
dynamic feedback loop
reward hacking
response discriminability
Innovation

Methods, ideas, or system contributions that make the work stand out.

EvoRS
self-evolving RL framework
on-policy experience
Reward-DAG
reward hacking
Weiyuan Li
Weiyuan Li
Alibaba Group
RLLLMAgent
Aili Chen
Aili Chen
Fudan University
Large Language ModelReasoning and PlanningLanguage AgentLLM Personalization
Xintao Wang
Xintao Wang
Fudan University
Role-Playing AgentsLanguage AgentsLarge Language ModelsKnowledge Graphs
Yikai Zhang
Yikai Zhang
Fudan university
Natural Language ProcessingAutonomous Agent
Q
Qingqing Dong
College of Cryptology and Cyber Science, Nankai University
J
Jinghan Xu
School of Data Science, Fudan University; Shanghai Key Laboratory of Data Science
H
Hongru Hou
School of Data Science, Fudan University; Shanghai Key Laboratory of Data Science
W
Wenxuan Zhao
Hello Group
C
Chengkun Lang
Hello Group
J
Jun Gao
Hello Group
Y
Yuanli Guo
Shanghai Key Laboratory of Data Science; School of Social Development and Public Policy, Fudan University
Hongcheng Guo
Hongcheng Guo
School of Data Science, Fudan University
LLMsMultimodal LLMs
Y
Yanghua Xiao
Shanghai Key Laboratory of Data Science; College of Computer Science and Artificial Intelligence, Fudan University
Deqing Yang
Deqing Yang
School of Data Science, Fudan University