Continual Reasoning Gym: Diagnosing and Harnessing Shared Reasoning in Continual RLVR

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过持续推理健身房环境诊断共享推理,并提出持续提示回放方法以解决持续强化学习中的性能下降问题。
📝 Abstract
Reinforcement learning with verifiable rewards (RLVR) commonly post-trains reasoning models on multiple tasks, while rerunning multitask RLVR (MTRL) as new tasks are added makes capability expansion costly. We therefore study continual RLVR, which updates the existing model as each task arrives. The central question is whether a model updated this way can perform as well as a jointly trained model. To answer this question, we introduce Continual Reasoning Gym, a continual-RLVR environment that organizes text and visual reasoning tasks into five task sequences. In this setting, we identify two key observations: Sequential RLVR exhibits modest forgetting, yet its final performance remains below that of MTRL. To understand the latter, we decompose final performance and show that forgetting accounts for only part of the gap. To explain the former, we identify shared reasoning: transferable reasoning structure allows training on one task to support others on average. We therefore introduce Continual Prompt Replay (CPR), which harnesses shared reasoning to improve learning on the arriving and future tasks by replaying previous-task prompts and regenerating their responses with the current policy. On average, only CPR reaches MTRL-level performance.
Problem

Research questions and friction points this paper is trying to address.

Continual RLVR
Model Updating
Performance Gap
Reasoning Transfer
Forgetting
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continual Reasoning
Shared Reasoning
Continual Prompt Replay
L
Lirui Luo
State Key Lab of General AI, School of Intelligence Science and Technology, Peking University; State Key Laboratory of General Artificial Intelligence, BIGAI
G
Guoxi Zhang
State Key Laboratory of General Artificial Intelligence, BIGAI
H
Hongming Xu
State Key Laboratory of General Artificial Intelligence, BIGAI
R
Rongqing Li
Beijing Institute of Technology
Cong Fang
Cong Fang
Peking University
machine learningoptmizationstatistics
Lifeng Fan
Lifeng Fan
University of California, Los Angeles
Artificial IntelligenceCognitive ModelingSocial Interaction