Last-Iterate Convergence of Policy Dynamics in Zero-Sum Networked Separable Markov Games

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了有限时域零和网络可分离马尔可夫博弈中直接策略更新方法的设计与分析不足的问题,通过提出熵正则乐观乘法权重更新(ER-OMWU)算法,实现近线性收敛速度求解近似纳什均衡。
📝 Abstract
Solving Nash equilibria for general multi-player Markov games is computationally intractable, while two-player zero-sum Markov games admit fast last-iterate policy-optimization methods. Finite-horizon zero-sum networked separable Markov games occupy an important middle ground: they retain global competition structure through pairwise interactions, while preserving computational tractability of Nash equilibria (NE) in the full-information and known-transition setting. Existing algorithms for this class either proceed through equilibrium-collapse arguments for a simplified setting where a single controller determines the transition probability, or backward dynamic programming that relies on equilibrium solvers at each stage. However, the design and analysis of direct policy-update approaches remain inadequate. To address this issue, we propose the entropy-regularized optimistic multiplicative weights update (ER-OMWU), a complementary single-loop policy dynamic that updates players'policies symmetrically and returns an approximate NE in the last iteration. We provide a first last-iterate convergence analysis of policy dynamics in the games of interest: after $\widetilde{O}(1/{\epsilon})$ iterations, the returned policy is an $\epsilon$-approximate Nash equilibrium. The result preserves the near-linear convergence rate achieved by policy optimization in two-player zero-sum Markov games, but extends the policy-dynamics viewpoint to a more complicated but structured multi-player setting.
Problem

Research questions and friction points this paper is trying to address.

Zero-Sum Networked Separable Markov Games
Policy Dynamics
Nash Equilibria
Direct Policy-Update
Innovation

Methods, ideas, or system contributions that make the work stand out.

Entropy-regularized optimistic multiplicative weights update (ER-OMWU)
Last-iterate convergence
Zero-sum networked separable Markov games
Approximate Nash equilibrium
💼 Related Jobs
No related jobs found.