Sparse Additive Off-Policy Evaluation for Reinforcement Learning with Potentially Limited Number of Trajectories

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种新的框架,通过稀疏加性结构的非线性函数类来解决强化学习中样本轨迹有限情况下的策略评估问题。
📝 Abstract
We develop a new framework for flexible, nonlinear, and interpretable off-policy evaluation for infinite-horizon reinforcement learning. To handle large state spaces and support transparent decision-making, we model the Q-function using a nonlinear function class with a sparse additive structure. We derive high-probability finite-sample error bounds for estimating the value function of a target policy and show that the bounds depend only logarithmically on the ambient dimension $d$, thereby alleviating the curse of dimensionality. In contrast to most existing theory for off-policy evaluation, which typically assumes access to many trajectories, our analysis guarantees accurate value estimation when either the number of trajectories or the time horizon is sufficiently large. In addition, we propose a group-sparsity-based feature screening procedure that identifies, with high probability, a reduced feature set containing all relevant covariates. Numerical experiments demonstrate the effectiveness of the proposed approach.
Problem

Research questions and friction points this paper is trying to address.

off-policy evaluation
reinforcement learning
sparse additive structure
limited number of trajectories
Innovation

Methods, ideas, or system contributions that make the work stand out.

sparse additive structure
off-policy evaluation
curse of dimensionality
group-sparsity-based feature screening
🔎 Similar Papers
No similar papers found.