Optimal Value Inference for Reinforcement Learning

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过近似最大贝尔曼算子并使用Neyman正交性提出去偏估计器,解决强化学习中的最优值离线推理问题。
📝 Abstract
We study offline inference for the optimal value in reinforcement learning. Two new nuisances are derived as fixed points of a self-induced Bellman equation, in which we approximate the maximum Bellman operator by its softmax correspondence. We propose a debiased estimator through the Neyman orthogonality and establish its asymptotic normality under diverging horizons even when the behavior policy changes with time, as long as the nuisances have the statistical rates that can be achieved by many machine learning methods. We provide a concrete estimating procedure for these nuisances and show they can lead to valid inference. Synthetic experiments validate the numerical performance of our inference method, and we implement it in real-life decision-making problems, including bike repositioning and AI agentic tool use.
Problem

Research questions and friction points this paper is trying to address.

Offline Inference
Optimal Value
Reinforcement Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

softmax approximation
Neyman orthogonality
debiased estimator
asymptotic normality
🔎 Similar Papers
No similar papers found.