Online Inference for Quantile Temporal Difference Learning in Distributional Reinforcement Learning

📅 2026-08-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the problem of online statistical inference in quantile temporal difference (QTD) learning. Under a generative model, the authors establish functional central limit theorems for both synchronous and asynchronous QTD, demonstrating that the averaged iterates weakly converge to a rescaled Brownian motion. Building on this asymptotic characterization, they propose a novel online inference method based on random scaling, which yields an asymptotically pivotal statistic without requiring storage of the full iteration trajectory. This approach substantially reduces memory overhead while maintaining theoretical rigor and computational efficiency, thereby enabling real-time statistical inference with low memory consumption and high practicality.
📝 Abstract
In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning. Assuming access to a generative model, we first establish functional central limit theorems for both synchronous and asynchronous QTD, which show that the averaged iterates of QTD converge weakly to a rescaled Brownian motion. We next provide online inference methods. Based on random scaling, the inference procedure constructs an asymptotically pivotal statistic for inference by using the information along the whole QTD path. Meanwhile, the proposed statistic can be computed online without storing the entire trajectory of QTD iterates. This substantially reduces the memory requirement and enables efficient statistical inference in distributional reinforcement learning.
Problem

Research questions and friction points this paper is trying to address.

quantile temporal difference learning
distributional reinforcement learning
online inference
statistical inference
functional central limit theorem
Innovation

Methods, ideas, or system contributions that make the work stand out.

quantile temporal difference
functional central limit theorem
online inference
random scaling
distributional reinforcement learning