Semi-Bandit Learning for Monotone Stochastic Optimization*
This paper addresses monotone stochastic optimization problems with unknown distributions—such as prophet inequalities and Pandora’s box—where classical approaches rely on full distributional knowledge. Method: We propose the first unified semi-bandit online learning framework that learns and approximates the optimal policy solely from observed samples of probed random variables, without any prior distributional information. Our approach integrates semi-bandit feedback modeling, a monotonicity-aware stochastic probing strategy, and refined regret analysis. Contribution/Results: We establish a near-optimal regret bound of $O(sqrt{T log T})$, which strictly improves upon fundamental lower bounds under both full-information and pure bandit settings for multiple canonical problems. This is the first result to break the long-standing dependence on distributional priors in stochastic optimization, providing both a general theoretical foundation and an efficient algorithmic pathway for online approximation in distribution-agnostic environments.