RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction
This work addresses the challenge that generative reward models, due to their comparative outputs, are incompatible with the scalar rewards required by reinforcement learning, thereby hindering effective training of large language models. To overcome this limitation, the paper proposes a Ranking-based Reward Construction (RRC) method, which innovatively introduces self-competitive ranking and anchor-guided ranking strategies to transform relative preference rankings into scalar reward signals suitable for reinforcement learning. By circumventing the constraints of conventional scalar reward formulation, RRC achieves substantially improved training performance on open-ended dialogue and reasoning benchmarks, consistently outperforming existing reward modeling approaches.