Constant Swap Regret in General-Sum Games via Optimistic Transition Matrices

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了在全信息反馈下有限多人常和博弈中实现恒定个体交换遗憾的学习动态方法,通过预测偏差收益并更新转移矩阵来达成。
📝 Abstract
We give deterministic and uncoupled learning dynamics for finite multiplayer general-sum games under full-information feedback that achieve constant individual swap regret, independent of the horizon $T$. With $n$ players and at most $m$ actions each, the individual swap regret of every player is $O(\sqrt{n} m \log m \log^{5/2}(nm))$ at every finite horizon. Each player predicts the deviation gains, then uses these predictions to update a row-stochastic transition matrix, and plays its stationary distribution. The proof combines a potential argument exploiting stationarity with a two-scale higher-order prediction analysis, using rooted-tree representations to handle the nonlinear dependence of deviation gains on the stationary distributions. An adversarially robust variant, obtained through a generic common-prefix switching wrapper, preserves the self-play bound up to a universal constant and guarantees individual swap regret at most $7\sqrt{m T \log m}$ in the adversarial setting.
Problem

Research questions and friction points this paper is trying to address.

swap regret
general-sum games
learning dynamics
Innovation

Methods, ideas, or system contributions that make the work stand out.

constant individual swap regret
optimistic transition matrices
full-information feedback
rooted-tree representations
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.