MAPPO-LCR: Multi-Agent Policy Optimization with Local Cooperation Reward in Spatial Public Goods Games

📅 2025-12-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
In spatial public goods games, strong payoff coupling, environmental non-stationarity, and population-level strategy dependencies pose significant challenges for existing approaches—such as evolutionary updating or independent reinforcement learning—which fail to capture inter-agent strategic interdependence. To address this, this paper introduces multi-agent proximal policy optimization (MAPPO) to the domain for the first time, proposing the MAPPO-LCR framework. Its key contributions are: (1) a centralized critic that explicitly models global policy coupling; (2) a local cooperation reward (LCR) mechanism, where reward signals are derived from neighborhood cooperation density to guide policy updates; and (3) unified decoupled execution with joint value estimation, preserving the original game structure. Experiments across diverse enhancement factors demonstrate that MAPPO-LCR consistently emergently fosters cooperation, substantially outperforming independent PPO in cooperation rate, convergence speed, and robustness.

Technology Category

Application Category

📝 Abstract
Spatial public goods games model collective dilemmas where individual payoffs depend on population-level strategy configurations. Most existing studies rely on evolutionary update rules or value-based reinforcement learning methods. These approaches struggle to represent payoff coupling and non-stationarity in large interacting populations. This work introduces Multi-Agent Proximal Policy Optimization (MAPPO) into spatial public goods games for the first time. In these games, individual returns are intrinsically coupled through overlapping group interactions. Proximal Policy Optimization (PPO) treats agents as independent learners and ignores this coupling during value estimation. MAPPO addresses this limitation through a centralized critic that evaluates joint strategy configurations. To study neighborhood-level cooperation signals under this framework, we propose MAPPO with Local Cooperation Reward, termed MAPPO-LCR. The local cooperation reward aligns policy updates with surrounding cooperative density without altering the original game structure. MAPPO-LCR preserves decentralized execution while enabling population-level value estimation during training. Extensive simulations demonstrate stable cooperation emergence and reliable convergence across enhancement factors. Statistical analyses further confirm the learning advantage of MAPPO over PPO in spatial public goods games.
Problem

Research questions and friction points this paper is trying to address.

Addresses payoff coupling in multi-agent spatial games
Introduces centralized critic for joint strategy evaluation
Enhances cooperation via local reward without altering game
Innovation

Methods, ideas, or system contributions that make the work stand out.

MAPPO introduces centralized critic for joint strategy evaluation
Local cooperation reward aligns policies with neighborhood cooperation density
Decentralized execution preserved while enabling population-level value estimation
🔎 Similar Papers
No similar papers found.
Z
Zhaoqilin Yang
State Key Laboratory of Public Big Data, College of Computer Science and Technology, Guizhou University, Guiyang, 550025, Guizhou, China
A
Axin Xiang
Institute of Cryptography and Data Security, Guizhou University, Guiyang, 550025, Guizhou, China
K
Kedi Yang
State Key Laboratory of Public Big Data, College of Big Data and Information Engineering, Guizhou University, Guiyang, 550025, Guizhou, China
T
Tianjun Liu
State Key Laboratory of Public Big Data, College of Computer Science and Technology, Guizhou University, Guiyang, 550025, Guizhou, China
Y
Youliang Tian
Institute of Cryptography and Data Security, Guizhou University, Guiyang, 550025, Guizhou, China