The Sample Complexity of Policy Learning with Mu-Resets

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work investigates the sample complexity of policy learning under the μ-reset interactive protocol, with a focus on how policy realizability influences sample efficiency. In this setting, the agent collects trajectories by resetting from an exploratory distribution μ. The authors formally establish, for the first time, the critical role of policy realizability in this framework and systematically analyze the dependence of sample complexity on the environment horizon H under varying concentrability assumptions, leveraging both all-policy and pushforward concentrability coefficients. Their main contributions include proving an exponential lower bound of exp(Ω(H)) on sample complexity under bounded all-policy concentrability, and establishing a tight characterization of exp(Θ(√H)) under bounded pushforward concentrability, thereby revealing a substantially improved sub-exponential dependence on H.
📝 Abstract
We study policy-based reinforcement learning under the $μ$-resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables the learner to sample trajectories from a given exploratory reset distribution $μ$, in addition to the starting distribution. We resolve the question raised by [KLS25] on the role of policy realizability for the sample complexity of this problem. Critically, the dependence on horizon $H$ is governed by the notion of coverage assumed of the reset distribution. Under bounded all-policy concentrability, we show a $\exp(Ω(H))$ sample complexity lower bound; with bounded pushforward concentrability, we show the dependence on horizon is tightly characterized as $\exp(Θ(\sqrt H))$.
Problem

Research questions and friction points this paper is trying to address.

sample complexity
policy learning
reinforcement learning
concentrability
horizon dependence
Innovation

Methods, ideas, or system contributions that make the work stand out.

sample complexity
policy learning
μ-resets
concentrability
horizon dependence
🔎 Similar Papers
No similar papers found.