A First-Order Learning Algorithm for Online Resource Allocation with Constant Regret

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
📝 Abstract
We study a finite-horizon online resource allocation problem with initial resource capacities proportional to the horizon. In each period, a request type is observed and one action is chosen from a finite menu. Each action earns a reward and consumes a vector of resources. The arrival types are independent and identically distributed, but their probabilities are unknown. We present a primal first-order learning policy that, in each period, performs one gradient ascent update of the action coordinates associated with the current request type. The policy achieves $O(1)$ expected additive regret relative to the hindsight optimum, with a bound independent of the horizon $T$. It does not solve any linear program, and the regret bound does not require a nondegeneracy assumption on the fluid linear program.
Problem

Research questions and friction points this paper is trying to address.

Online Resource Allocation
Constant Regret
Finite-horizon
Innovation

Methods, ideas, or system contributions that make the work stand out.

First-Order Learning
Online Resource Allocation
Constant Regret
Gradient Ascent
Nondegeneracy Assumption
🔎 Similar Papers
M
Menglong Li
Department of Decision Analytics and Operations, College of Business, City University of Hong Kong, Hong Kong
Jiawei Zhang
Jiawei Zhang
Professor of Operations Management, New York University
Operations ResearchOperations Management