A Convergence Framework for Deep $V$-Learning: Error Propagation and Sharp Action-Gap Bounds

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文解决了深度V学习中的收敛性问题,通过分解六种残差并分析其L^p范数来控制策略损失,提出了一种新的误差传播框架。
📝 Abstract
We establish convergence bounds for deep $V$-learning with horizon $H$. The algorithm fits a scalar value function to targets from executed transitions and selects actions using a predictive model and the value function. For current observed-successor targets with fresh true-kernel outcomes, the conditional mean is $\mathcal{T}^βV$, which averages over behavior-policy actions. The Bellman optimality update is $\mathcal{T} V$. We decompose the update error into six residuals: fitting, transition reuse, target construction, replay, action selection, and exploration. Under $L^s$ concentrability, their $L^p$ norms ($p=s/(s-1)$) control expected $L^1$ policy loss. The bound explicitly weights residuals from only the last $H-1$ update blocks, plus an initialization term for shorter runs. We quantify the cost of a shared sampling distribution across horizon levels. For statistical error bounds of order $n^{-ν}$, we derive optimal continuous allocations and an integer allocation whose objective is within a factor $2^ν$ of the constrained optimum. A margin condition with exponent $α$ gives action error of order $Λ^{1+α/p}$, where $Λ$ combines network drift and score error; a one-step construction proves the exponent sharp. Bounds on the distance between frozen and optimal scores transfer an optimal-gap condition to frozen-iterate gap bounds while retaining the mass of optimal ties. Survival probabilities and coverage conditions at deployment yield bounds for policies selected with approximate scores. Separate spatial ReLU networks per horizon level give a conditional neural regression rate, and the finite-state case gives a log-free expected fit rate. These results give expected policy-loss consistency for the fixed-horizon generative-reset approximate-ERM procedure with exact action scores and provide an explicit residual-decay criterion for FIFO/interleaved SGD.
Problem

Research questions and friction points this paper is trying to address.

convergence bounds
deep V-learning
error propagation
action-gap bounds
policy loss
Innovation

Methods, ideas, or system contributions that make the work stand out.

Convergence Bounds
Error Decomposition
Residuals Control
Optimal Allocation
Action-Gap Bound
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yury Kolomeytsev
Faculty of Computational Mathematics and Cybernetics, Lomonosov Moscow State University