Toward Optimal Switching Regret for Multi-Armed Bandits with Oblivious Adversary

📅 2026-09-11
📈 Citations: 0
✹ Influential: 0
📄 PDF
🀖 AI Summary
本文解决了圚对抗性倚臂老虎机䞭圓变化次数S未知时劂䜕蟟到最䌘切换遗憟的问题通过结合固定仜额孊习噚和二进制闎隔子皋序的方法。
📝 Abstract
We study switching regret in adversarial multi-armed bandits, where the learner competes with an arm sequence that changes at most $S$ times. When $S$ is known, an optimal expected regret of $\widetilde{\mathcal{O}}(\sqrt{(S+1)KT})$ is obtainable [Auer et al., 2002]. However, when $S$ is unknown, Marinov and Zimmert [2021] show that this guarantee is impossible under an adaptive adversary. In this paper, we show that a single algorithm achieves $\widetilde{\mathcal{O}}(\sqrt{(S+1)KT})$ expected regret for every $S$ against an oblivious adversary, resolving an open problem of Auer et al. [2019b]. Our algorithm combines a fixed-share learner initialized with a small learning rate and dyadic-interval subroutines that search for local improvements using randomized learning rates and implicit exploration. Importantly, a non-uniform prior favors following the main learner, keeping the cost of maintaining many subroutines small. When the subroutines accumulate sufficient improvement over the main learner, its learning rate doubles, allowing adaptation to the unknown number of comparator switches $S$.
Problem

Research questions and friction points this paper is trying to address.

adversarial multi-armed bandits
oblivious adversary
switching regret
unknown S
Innovation

Methods, ideas, or system contributions that make the work stand out.

oblivious adversary
switching regret
adaptive learning rate
🔎 Similar Papers
No similar papers found.