Optimal Alternating Regret for Online Learning and Games

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文解决了在线线性优化和凸优化中的最优交替遗憾问题,通过提出新算法实现了对数级改进,并在二人博弈中达到更快的均衡收敛速度。
📝 Abstract
We settle the minimax-optimal alternating regret, a regret notion motivated by alternating learning dynamics in games, for both online linear optimization (OLO) and online convex optimization (OCO). For OLO over the probability simplex $Δ_d$, we give an algorithm with $O(\log d)$ alternating regret that remains a constant for any time horizon $T$, and a matching lower bound. Our constant regret bound significantly improves previous results with $O(\log ^{2/3}d \cdot T^{1/3})$ regret [Cevher, Cutkosky, Kavis, Piliouras, Skoulakis, Viano, NeurIPS 2023, Hait, Li, Luo, Zhang, COLT 2025]. As a result, we obtain alternating learning dynamics with $O(\log d /T)$ convergence to Nash equilibria in two-player zero-sum games and $O(\log d /T)$ convergence to coarse correlated equilibria in two-player general-sum games. This is the first uncoupled learning dynamics with $O(1/T)$ convergence to CCE in two-player general-sum games, while all prior works suffer additional $\log T$ factors. For general OCO over a $d$-dimensional compact convex set, we give an algorithm with $O(d\log (1+T/d))$ alternating regret, improving the previous best of $\widetilde{O}(d^{2/3}T^{1/3})$. We also prove a matching lower bound of $Ω(d\log (1+T/d))$, showing that the $Ω(\log T)$ factor is unavoidable.
Problem

Research questions and friction points this paper is trying to address.

alternating regret
online linear optimization
online convex optimization
Nash equilibria
coarse correlated equilibria
Innovation

Methods, ideas, or system contributions that make the work stand out.

Optimal Alternating Regret
Online Linear Optimization
Online Convex Optimization
Nash Equilibrium Convergence
Coarse Correlated Equilibria
🔎 Similar Papers
No similar papers found.