Lipschitz Bandits with Arbitrary Feedback Delays

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses Lipschitz multi-armed bandits with continuous action spaces under arbitrary feedback delays. For both stochastic and adversarial settings, we propose algorithms based on elimination strategies and EXP3, introducing a scaling dimension to precisely characterize the impact of delays. Theoretically, this work provides the first characterization of the additional cost induced by arbitrary delays, deriving a regret bound of Õ(T^((dz+1)/(dz+2)) + √D). This result matches optimal delay-free performance while achieving tight bounds on delay-sensitive terms, effectively resolving continuous decision-making challenges in delayed environments.
📝 Abstract
The Lipschitz bandit problem extends the traditional multi-armed bandit framework to continuous action spaces by assuming that the reward functions satisfy a Lipschitz condition. This work investigates Lipschitz bandits under arbitrary feedback delays, where reward signals are not received immediately upon taking an action but after an arbitrarily chosen delay. We consider both stochastic and adversarial reward settings, proposing an elimination-based algorithm and an EXP3-based algorithm, respectively. For both settings, our algorithms achieve a regret bound of $\tilde{O}\left(T^{\frac{d_z+1}{d_z+2}}+\sqrt{D}\right)$ over a time horizon $T$ with total delay $D$, where the main difference between settings lies in the definition of the zooming dimension $d_z$. Our bounds match existing delay-free regret guarantees for Lipschitz bandits and characterize the additional $\tilde{O}(\sqrt{D})$ impact introduced by feedback delays.
Problem

Research questions and friction points this paper is trying to address.

Lipschitz bandits
arbitrary feedback delays
continuous action spaces
stochastic rewards
adversarial rewards
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lipschitz Bandits
Arbitrary Feedback Delays
Regret Bound
Zooming Dimension
Continuous Action Spaces
🔎 Similar Papers
2024-07-24arXiv.orgCitations: 4