A Modern Introduction to Online Learning
This work addresses worst-case online optimization, unifying the study of online convex and non-convex optimization over Euclidean and non-Euclidean domains—including simplices and matrix manifolds—under a regret minimization framework. We propose a parameter-free, adaptive algorithmic framework that supports unbounded decision sets and unknown gradient magnitudes. Unifying online mirror descent (OMD) and follow-the-regularized-leader (FTRL), we reformulate first- and second-order methods and, for the first time, integrate convex surrogate losses, randomization schemes, and multi-armed bandit feedback—both adversarial and stochastic—into this coherent paradigm. Our theoretical analysis is self-contained, elementary, and accessible without prerequisites; all algorithms achieve tight, optimal regret bounds. The resulting framework establishes a universal, concise, and pedagogically transparent foundation for modern online learning, substantially lowering both theoretical barriers and practical implementation complexity.