Scheduling Servers with Stochastic Bilinear Rewards
This paper addresses online scheduling in a parallel queue system with multiple job classes and multiple servers, where rewards are unknown, dynamically stochastic, and exhibit a bilinear structure. The objective is to jointly maximize cumulative reward and minimize job holding delay (i.e., holding cost), while ensuring system stability—namely, throughput optimality and bounded queue lengths. We propose the first distributed algorithm integrating three key components: (i) dynamic learning of bilinear bandit rewards, (ii) weighted proportional-fair scheduling, and (iii) marginal-cost correction. Theoretically, the algorithm achieves a sublinear regret bound and guarantees bounded expected queue lengths. Empirically, it significantly outperforms existing baselines in both cumulative reward and average delay across computational service and online platform scenarios.