Diffusion Models Meet Contextual Bandits with Large Action Spaces

📅 2024-02-15

🏛️ arXiv.org

📈 Citations: 5

✨ Influential: 0

career value

213K/year

🤖 AI Summary

To address inefficient exploration in contextual bandits with large action spaces—where the absence of prior knowledge incurs high statistical and computational costs—this paper introduces diffusion Thompson Sampling (dTS), the first method to incorporate pretrained diffusion models into this setting. dTS leverages diffusion models to implicitly capture reward correlations among actions, thereby constructing a generative Bayesian prior that enables efficient posterior sampling and exploration. We establish a sublinear regret bound for dTS, proving its theoretical soundness. Empirically, dTS significantly outperforms classical baselines—including LinUCB, NeuralUCB, and standard Thompson Sampling—across multiple large-scale action benchmarks, demonstrating both statistical efficacy and computational feasibility. Our core contribution is the pioneering integration of diffusion modeling with Bayesian online decision-making, yielding a scalable, data-efficient exploration paradigm for high-dimensional action spaces.

Technology Category

Application Category

📝 Abstract

Efficient exploration is a key challenge in contextual bandits due to the large size of their action space, where uninformed exploration can result in computational and statistical inefficiencies. Fortunately, the rewards of actions are often correlated and this can be leveraged to explore them efficiently. In this work, we capture such correlations using pre-trained diffusion models; upon which we design diffusion Thompson sampling (dTS). Both theoretical and algorithmic foundations are developed for dTS, and empirical evaluation also shows its favorable performance.

Problem

Research questions and friction points this paper is trying to address.

Efficient decision-making in large contextual bandit action spaces

Leveraging diffusion models as priors for complex action distributions

Developing practical posterior approximation algorithms for flexible strategies

Innovation

Methods, ideas, or system contributions that make the work stand out.

Leverages pre-trained diffusion models as priors

Develops diffusion-based decision framework for bandits

Efficiently approximates posteriors under diffusion priors

🔎 Similar Papers

Bellman Diffusion Models