An Evolutionary Computation Framework for Multi-Agent Q-Learning with Mean-Field Environmental Feedback

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过结合均场环境反馈的多智能体Q学习框架,解决了网络化群体中个体适应、局部互动与环境变化间的相互作用问题。
📝 Abstract
Multi-agent reinforcement learning in networked populations is governed by the interaction between individual adaptation, local encounters, and changing environmental conditions. To study this interaction, we formulate a coupled learning--environment model in which agents update stateless $Q$-values on a fixed graph, while their population-average behavior drives an environmental variable that dynamically modifies the payoff matrix. Under a first-order mean-field closure, we derive a deterministic transport equation for the population distribution of $Q$-values and couple it with a projected discrete update for the environmental state. The resulting model is evaluated against finite-network Monte Carlo simulations on random regular, Erdős--Rényi, Barabási--Albert, and random geometric graphs. Across the tested parameter ranges, the mean-field system reproduces the main macroscopic cooperation and environmental trajectories, and the trajectory-level root-mean-square error generally decreases with population size and average degree. The analysis further shows that environmental feedback reshapes the learned action-value ordering, while reinforcing feedback can produce pronounced dependence on the initial learning bias and resource level. The environmental timescale also plays an important role: a rapid response can drive the resource state to a boundary before learning adapts, whereas a slower response preserves the interaction between behavioral learning and environmental recovery. These results provide a population-level description of coupled reinforcement learning and environmental dynamics and characterize the performance of the mean-field approximation within the tested network and parameter ranges.
Problem

Research questions and friction points this paper is trying to address.

Multi-agent reinforcement learning
Environmental feedback
Mean-field approximation
Networked populations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evolutionary Computation
Multi-Agent Q-Learning
Mean-Field Environmental Feedback
Deterministic Transport Equation
Reinforcement Learning