A Finite Sample Analysis for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

📅 2026-08-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过分析同步分位数时差学习在分布强化学习中的全局有限样本保证,解决了其收敛性问题,利用了奖励累积分布函数的顺序单调性和分布贝尔曼算子的W_∞收缩。
📝 Abstract
We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms. A global comparison argument, based on the order monotonicity of reward cumulative distribution functions and the $W_\infty$ contraction of the distributional Bellman operator, brings an arbitrarily initialized iterate into a local neighborhood. Inside that neighborhood, we linearize the QTD mean field. Its Jacobian is a nonsingular $M$-matrix, and the associated positive semigroup permits a variance-sensitive martingale analysis. For stepsizes $α_t=c(t+1)^{-a}$ with $a\in(1/2,1)$, the leading last-iterate fluctuation is of order $\widetilde O\bigl(T^{-a/2}/\sqrt{1-γ}\bigr)$ and has no polynomial dependence on the number of quantiles. The deterministic transient and the required burn-in can still depend on the smallest Bellman-target density, which is of order $m^{-1}$ in the worst case. The result therefore distinguishes sharply between the local stochastic fluctuation and the global sample complexity.
Problem

Research questions and friction points this paper is trying to address.

Quantile Temporal Difference Learning
Distributional Reinforcement Learning
Finite Sample Analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Quantile Temporal Difference Learning
Distributional Reinforcement Learning
Finite Sample Analysis
Variance-Sensitive Martingale Analysis
Global and Local Stability Mechanisms