From Memorization to Absorption: Mixed-Policy RL for Continual Knowledge Injection

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大型语言模型持续知识注入中的泛化问题,提出Golden-GRPO Injection方法,通过混合策略强化学习实现优于监督微调的知识吸收效果。
📝 Abstract
Continual knowledge injection is essential for keeping large language models up-to-date in a fast-evolving world. Existing methods rely on supervised fine-tuning (SFT), which memorizes injected facts in their training format but fails to generalize across paraphrasing, document combinations, and reasoning. To address this, we propose Golden-GRPO Injection (GRIN), a three-stage self-learning framework for continual knowledge injection. Golden-GRPO is a mixed-policy reinforcement learning algorithm designed specifically for knowledge injection, which injects a golden answer to provide learning signal even when on-policy rollouts fail on novel facts. We further introduce Blank and Counter, two document-level benchmarks targeting novel acquisition and counterfactual overwrite respectively, each evaluating single-fact recall, multi-source retrieval, and inferential reasoning. Our experiments establish a clear empirical claim: mixed-policy reinforcement learning enables knowledge absorption beyond what supervised fine-tuning can achieve. GRIN substantially outperforms SFT and mixed-policy RL baselines on the harder question types while matching them on basic fact recall.
Problem

Research questions and friction points this paper is trying to address.

Continual Knowledge Injection
Supervised Fine-Tuning
Generalization
Paraphrasing
Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

mixed-policy reinforcement learning
continual knowledge injection
self-learning framework
Golden-GRPO
🔎 Similar Papers
No similar papers found.