Bandits with Probing: Optimal Regret and the Limits of Winner Feedback

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了多臂赌博机问题中通过探测最多k个臂来优化后悔界的问题,使用了赢家反馈机制,并确定了两种极小极大法则。
📝 Abstract
A learner probes at most $k$ of $n$ arms each round, receives the maximum of their rewards in $[0,1]$, and competes with the best fixed arm. When does the probing advantage pay for learning? We determine two minimax laws. Under independent stochastic rewards with winner feedback (the maximum and a winning label), or on arbitrary fixed sequences given a single signed contrast between block maxima, the minimax regret has order $Φ_{n,k}(T)=\min\{\frac{n-k}{n}T,\frac{n-k}{k}\}$, $2\le k<n$. Under winner feedback, both arbitrary joint i.i.d. rewards and fixed sequences have minimax regret of order $R_{n,k}(T)=\frac{n-k}{n}\min\{T,\frac{n+T}{k},\sqrt{\frac{nT}{k}}\}$. Both laws have universal constants and anytime upper bounds. The first reduces regret to a pure coverage cost: same-round contrasts absorb the stability cost, and independence permits exact resampling whose gains fund sample advancement. The second adds a learning cost that becomes comparable to coverage at horizon $n$; beyond $nk$, numerical maxima improve over labels alone. The lower bound allows every adaptive action size.
Problem

Research questions and friction points this paper is trying to address.

bandits
probing
regret
winner feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bandits with Probing
Winner Feedback
Minimax Regret
Independent Stochastic Rewards
Arbitrary Fixed Sequences
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yongjie Guan
Zhejiang University of Technology