π€ AI Summary
This work addresses the instability in multi-agent coordination caused by supervisor belief estimation errors and false alarms arising from Byzantine agentsβ joint actions. To tackle this challenge, the paper proposes a Distributed Team Coordination Algorithm (DTOA), which introduces a supervisor network into zero-sum potential team games for the first time. DTOA integrates team fictitious play with supervisor-based distributed belief learning and incorporates a Byzantine-resilient mechanism. Theoretical analysis establishes that the belief estimates converge to a team Nash equilibrium, provides an asymptotic bound on the honest teamβs Nash gap, and offers probabilistic guarantees for identifying Byzantine agents. Experimental results demonstrate that DTOA significantly outperforms baseline methods in Markov decision environments.
π Abstract
In this paper, we study zero-sum potential team games with a supervisor network, where agents rely on supervisor-provided belief information rather than accurate common beliefs. The main challenge is that such belief information can be inaccurate because of supervisors' belief-estimation errors and the misreporting of joint actions by Byzantine teams. We propose the distributed team-orchestrating algorithm (DTOA), which combines team fictitious play with supervisor-based distributed belief learning. We prove the convergence of supervisors' belief estimates and establish that the induced learning dynamics converge to a near team-Nash equilibrium (TNE) in terms of the team-Nash gap (TNG). In the Byzantine setting, we consider a misreporting attack model and develop a Byzantine-resilient DTOA. We further provide probabilistic guarantees for Byzantine-team identification and establish an asymptotic bound on the honest TNG. Numerical experiments illustrate the theoretical findings, compare DTOA with baseline learning methods, and evaluate its performance in a Markov decision process setting.