Offline RL with Smooth OOD Generalization in Convex Hull and its Neighborhood

📅 2025-06-10
🏛️ International Conference on Learning Representations
📈 Citations: 1
Influential: 0
📄 PDF
🤖 AI Summary
Offline reinforcement learning suffers from overestimation of Q-values for out-of-distribution (OOD) actions due to distributional shift, while existing constraint-based methods are overly conservative, impairing generalization and policy optimization. To address this, we propose the Convex-Hull Neighborhood (CHN) safe generalization framework. First, we formally define the CHN region to strike a principled balance between OOD exploration and in-distribution reliability. Second, we design the Smooth Bellman Operator (SBO), which—uniquely—establishes theoretical approximability guarantees for Q-values of OOD actions within the CHN. Third, we introduce neighborhood-weighted Q-function smoothing coupled with offline policy co-optimization. Evaluated on the D4RL benchmark, our method significantly outperforms state-of-the-art approaches: it yields more accurate Q-value estimation, superior policy performance, and improved computational efficiency.

Technology Category

Application Category

📝 Abstract
Offline Reinforcement Learning (RL) struggles with distributional shifts, leading to the $Q$-value overestimation for out-of-distribution (OOD) actions. Existing methods address this issue by imposing constraints; however, they often become overly conservative when evaluating OOD regions, which constrains the $Q$-function generalization. This over-constraint issue results in poor $Q$-value estimation and hinders policy improvement. In this paper, we introduce a novel approach to achieve better $Q$-value estimation by enhancing $Q$-function generalization in OOD regions within Convex Hull and its Neighborhood (CHN). Under the safety generalization guarantees of the CHN, we propose the Smooth Bellman Operator (SBO), which updates OOD $Q$-values by smoothing them with neighboring in-sample $Q$-values. We theoretically show that SBO approximates true $Q$-values for both in-sample and OOD actions within the CHN. Our practical algorithm, Smooth Q-function OOD Generalization (SQOG), empirically alleviates the over-constraint issue, achieving near-accurate $Q$-value estimation. On the D4RL benchmarks, SQOG outperforms existing state-of-the-art methods in both performance and computational efficiency.
Problem

Research questions and friction points this paper is trying to address.

Addresses Q-value overestimation in offline RL for OOD actions
Enhances Q-function generalization in convex hull and neighborhood regions
Reduces over-constraint issues to improve policy performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Enhances Q-function generalization in OOD regions
Introduces Smooth Bellman Operator for Q-value updates
Achieves near-accurate Q-value estimation efficiently
🔎 Similar Papers
Q
Qingmao Yao
School of Mathematical Sciences, Beihang University
Z
Zhichao Lei
School of Physics, Beihang University
Tianyuan Chen
Tianyuan Chen
Beihang University
reinforcement learning
Z
Ziyue Yuan
School of Mathematical Sciences, Beihang University
X
Xuefan Chen
School of Mathematical Sciences, Beihang University
J
Jianxiang Liu
School of Artificial Intelligence, Beihang University; Key Laboratory of Mathematics, Informatics and Behavioral Semantics, MoE, Beihang University
F
Faguo Wu
School of Artificial Intelligence, Beihang University; Key Laboratory of Mathematics, Informatics and Behavioral Semantics, MoE, Beihang University; Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University; Zhongguancun Laboratory
X
Xiao Zhang
School of Mathematical Sciences, Beihang University; Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing, Beihang University; Hangzhou International Innovation Institute of Beihang University