Multi-step Proximal Policy Improvement in Offline Reinforcement Learning

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出多步近端策略改进(MPI)方法,解决离线强化学习中策略更新需保持价值估计可靠同时超越行为分布的问题。
📝 Abstract
Offline reinforcement learning (RL) must reconcile two competing requirements: policy updates should stay near dataset-supported actions to keep value estimates reliable, yet meaningful gains often require moving beyond the behavior distribution. We develop a geometric view of offline actor updates by modeling policies as a probability manifold endowed with a chosen metric geometry. Under this lens, a broad class of offline actor objectives can be interpreted as a single proximal policy improvement step (SPI), i.e., an implicit discretization of a manifold gradient flow induced by a critic-defined energy. Building on this insight, we propose multi-step proximal policy improvement (MPI), a plug-in refinement mechanism that composes sequential re-centered proximal steps. MPI enables controlled policy improvement beyond dataset support while retaining proximal control at each refinement. The framework accommodates multiple policy geometries and admits practical instantiations for deterministic and diagonal-Gaussian policies. Experiments on D4RL benchmarks show that small numbers of MPI refinements improve strong offline baselines, including TD3+BC, ReBRAC, and IQL, on many tasks. Focused diagnostics further distinguish re-centered refinement from fixed-objective update scheduling and characterize limitations under critic error.
Problem

Research questions and friction points this paper is trying to address.

offline reinforcement learning
policy improvement
dataset support
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-step proximal policy improvement
offline reinforcement learning
probability manifold
policy geometry
refinement mechanism
🔎 Similar Papers
S
Soohyun Choi
Department of Electronic Engineering, Hanyang University, Seoul, Republic of Korea
S
Seonvin Cho
Department of Electronic Engineering, Hanyang University, Seoul, Republic of Korea
Songnam Hong
Songnam Hong
Hanyang University
Machine LearningInformation TheoryOptimization