Continuity-Free Near-Minimax Leading-Order Regret for CVaR-UCBVI

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文解决了CVaR强化学习中的遗憾界问题,通过Bernstein CVaR-UCBVI算法,在无需连续性假设下达到近似最优的遗憾率。
📝 Abstract
For finite-horizon tabular CVaR reinforcement learning, prior work proves a $\widetilde{O}(τ^{-1}\sqrt{SAK})$ leading regret bound for arbitrary normalized return laws and the sharper $\widetilde{O}(\sqrt{SAK/τ})$ rate under a density lower bound. We show that the same Bernstein CVaR-UCBVI algorithm attains the sharper rate without continuity assumptions. The key is a selected-budget self-bound: the conditional variance of the episode shortfall is at most $τ$ plus the value-estimation width. Substitution into the original Bernstein decomposition yields, with high probability, $\widetilde{O}(\sqrt{SAK/τ}+(SAHK^{1/4}+S^2AH)/τ)$ regret for arbitrary normalized return laws, including atomic, mixed, and continuous laws. The $τ^{-1/2}$ leading term matches the expected-regret minimax lower bound up to logarithmic factors. Thus Bernstein CVaR-UCBVI is minimax-optimal over the full return-law class in the leading-order regime; the lower-order terms retain their $τ^{-1}$ dependence.
Problem

Research questions and friction points this paper is trying to address.

CVaR-UCBVI
finite-horizon
tabular reinforcement learning
regret bound
continuity assumptions
Innovation

Methods, ideas, or system contributions that make the work stand out.

CVaR-UCBVI
selected-budget self-bound
minimax-optimal
regret bound
continuity-free
🔎 Similar Papers
2024-02-27IEEE Transactions on Information TheoryCitations: 1
2024-02-05arXiv.orgCitations: 1
💼 Related Jobs
No related jobs found.
Y
Yuanlong Chen
MiniMax