Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决大语言模型策略优化中的稳定性-探索困境,提出环境正则化策略优化(ERPO),通过引入Query-KL项来控制查询分布漂移,保持探索能力的同时提高准确性和稳定性。
📝 Abstract
Policy optimization (PO) for Large Language Models faces a stability--exploration trade-off, currently mediated by an action-side Policy-KL regularizer. This puts practitioners in a double bind: keeping Policy-KL constrains response behavior and consumes the action-side exploration budget, while dropping it leaves the optimization without an explicit drift control. We argue for an alternative that breaks the dilemma by moving regularization to the input side. As training progresses, the distribution over training queries induced by the current policy drifts unchecked from its pre-RL reference distribution. Concretely, Environment-Regularized Policy Optimization (ERPO) introduces a Query-KL (QKL) term that bounds this query distribution shift, together with a dataset-static reference-derived per-query weight that biases each per-query update toward queries typical under the reference. The QKL gradient flows strictly through the query likelihood; the response score function used by policy-gradient estimators does not appear in the QKL term, so QKL exerts no direct gradient pressure on the response distribution---exploration is preserved. ERPO plugs into GRPO/PPO/REINFORCE-style pipelines without additional forward passes. On six mathematical reasoning benchmarks, ERPO replaces the standard Policy-KL regularizer while achieving effective control over query distribution drift, delivering stronger accuracy and substantially more stable behavior under high-temperature decoding and long-horizon training.Our source code are available at https://github.com/alibaba/ERPO
Problem

Research questions and friction points this paper is trying to address.

Policy Optimization
Large Language Models
Stability-Exploration Trade-off
Policy-KL Regularizer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Environment-Regularized Policy Optimization
Query-KL
input-side regularization
policy optimization for LLMs
query distribution drift control
🔎 Similar Papers
2024-07-09Neural Information Processing SystemsCitations: 3
X
Xianlei Zhou
AMAP, Alibaba Group
X
Xiangdi Meng
AMAP, Alibaba Group
Y
Yu He
Xi’an Jiaotong University
T
Tianyu Qi
JD.com
S
Shuyan Guan
AMAP, Alibaba Group
X
Xianli Zhang
AMAP, Alibaba Group
J
Jian Zhang
Beijing Normal University
Xin Li
Xin Li
Alibaba Group
natural language processing
Qika Lin
Qika Lin
National University of Singapore | NTU | XJTU | BIT
Knowledge ReasoningNeurosymbolic AIMulti-modalRobustness & SecurityAI for Healthcare
J
Jun Liu
Xi’an Jiaotong University