When Greedy Sampling Explores: KL-Regularized Contextual Bandits without Eluder-Dimension Dependence

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过KL正则化方法在奖励和偏好反馈下,使用贪婪采样解决上下文老虎机问题,实现了与eluder维度无关的对数后悔界。
📝 Abstract
We study KL-regularized contextual bandits under both reward and preference feedback. We show that greedy sampling can achieve logarithmic regret without explicit dependence on the eluder dimension. For reward feedback, we establish an eluder-dimension-independent regret bound for a simple greedy algorithm that directly samples from the Gibbs policy induced by the estimated reward. We further extend this result to preference feedback under both the general preference and Bradley--Terry models, while also sharpening existing dimension-dependent guarantees. Our analysis reveals a trade-off between greedy sampling and upper confidence bound-style exploration: greedy sampling enjoys stronger guarantees when KL regularization is sufficiently strong, whereas additional exploration becomes preferable as the regularization weakens.
Problem

Research questions and friction points this paper is trying to address.

contextual bandits
KL-regularization
greedy sampling
regret bound
eluder dimension
Innovation

Methods, ideas, or system contributions that make the work stand out.

Greedy Sampling
KL-Regularization
Contextual Bandits
Eluder Dimension Independence
Logarithmic Regret
🔎 Similar Papers
Z
Zichen Wang
Department of ECE and CSL, University of Illinois Urbana-Champaign
H
Haoyang Hong
School of Electrical Engineering and Computer Science, Oregon State University
Huazheng Wang
Huazheng Wang
Assistant Professor, Oregon State University
Reinforcement LearningMachine LearningInformation Retrieval