Posterior Tempering Explains Variance Inflation in Linear and Generalized Linear Thompson Sampling

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文研究了α-TS算法解决随机广义线性bandit问题,通过引入分数后验来解释方差膨胀现象,并在一般条件下得到了最优的遗憾界。
📝 Abstract
We study a variant of the Thompson Sampling (TS) algorithm, called $α$-TS, for solving stochastic generalized linear bandit problems. Existing analyses of TS require inflating the posterior variance to derive near-optimal regret guarantees. We formalize the idea of variance inflation by introducing $α$-TS that uses a fractional or $α$-posterior instead of the standard posterior. Our main contribution is to identify general regularity conditions on the prior and reward distributions that enable a regret analysis of $α$-TS without assuming any tractable approximation of the posterior distribution, unlike previous works. For a specific choice of $α\propto d^{-1}$, our general regret bound yields the best known regret bound of $O(d^{3/2}\sqrt{T}\log T)$ for both the exponential and sub-Gaussian families of reward distributions. We further provide an $α$-dependent lower bound showing that the regret constant depends on the product $αd$, and that when $α\propto d^{-1}$ the regret scales as $Ω(d^{3/2}\sqrt{T})$, explaining the origin of the $d^{3/2}$ factor in the upper bound. Our proof technique adapts and combines recent advancements in the analysis of linear bandit problems with first- and second-order posterior concentration theory from the Bayesian statistics literature.
Problem

Research questions and friction points this paper is trying to address.

Thompson Sampling
Variance Inflation
Generalized Linear Bandits
Innovation

Methods, ideas, or system contributions that make the work stand out.

α-TS
Variance Inflation
Regret Analysis
Posterior Concentration
Generalized Linear Bandits
🔎 Similar Papers