What preferences can - and cannot - predict in multi-agent online learning

πŸ“… 2026-08-13
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of preference prediction dynamics in multi-agent online learning by integrating game theory, Follow-The-Regularized-Leader (FTRL) algorithms, and dynamical systems theory. We propose an aggregation bias resilience criterion to bridge existing theoretical gaps. The research elucidates the equivalence boundary between preference stability and dynamic stability, proving the asymptotic stability of preference representations under subgames while disproving general equivalence through counterexamples. Furthermore, we establish sufficient conditions guaranteeing asymptotic stability for arbitrary pure strategy sets. These findings clarify the predictive capacity of game preference graphs regarding no-regret learning dynamics, thereby refining the theoretical framework for understanding convergence in multi-agent systems.
πŸ“ Abstract
We examine the interplay between ordinal, preference-based solution concepts in games and the long-run behavior of game dynamics, asking in particular to what extent the combinatorial data of a game -- its preference graph -- determine the outcomes of no-regret learning dynamics -- such as follow-the-regularized-leader (FTRL). In one direction, we show that the skeleton of every dynamically stable set (i.e. the set of pure profiles it contains) must also be preferentially stable, that is, it must be closed under profitable deviations. We then ask the converse question: when do preferences determine the long-run behavior of the players' learning dynamics? We begin by showing that preferences characterize asymptotic stability in the case of subgames -- i.e. subsets of pure profiles obtained by restricting players' action sets. Beyond this case however, the equivalence between dynamic and preferential stability collapses: concretely, we construct a three-player game with a preferentially stable set whose span is dynamically unstable, showing in this way that preferences do not suffice as a criterion of dynamic stability. We then bridge this gap via the notion of resilience under aggregate deviations, an easy-to-check payoff-based condition that guarantees asymptotic stability of arbitrary spans of pure strategies.
Problem

Research questions and friction points this paper is trying to address.

multi-agent online learning
no-regret learning dynamics
preference graph
dynamic stability
preferential stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

No-regret learning
Preference graph
Dynamic stability
Resilience under aggregate deviations
FTRL
πŸ”Ž Similar Papers
No similar papers found.