Learning under Opponent Unawareness in Linear-Quadratic Stochastic Games

๐Ÿ“… 2026-08-08
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses decentralized learning in infinite-horizon non-zero-sum linear-quadratic stochastic games under minimal information conditions, where agents have no knowledge of opponentsโ€™ strategies and can only observe the public state and their own actions. The work proposes an asynchronous ฮต-greedy iterative least-squares algorithm that enables each agent to learn independently without identifying system parameters. For the first time, it is theoretically established that decentralized learning under such minimal information structures converges almost surely to the full-information Nash equilibrium, with explicit characterization of the convergence rate. Numerical experiments demonstrate that limited information disclosure reduces firm profits and social welfare while intensifying market concentration, whereas public disclosure of aggregate output significantly accelerates convergence and mitigates welfare losses.
๐Ÿ“ Abstract
As firms increasingly deploy machine learning for strategic decision-making, understanding algorithmic interactions has become central to operations research and economics. This paper studies learning in infinite-horizon, nonzero-sum linear-quadratic stochastic games under a radically uncoupled information structure, where players are either unaware of opponents or strategically oblivious, observing only a common state and their own action history. Under this minimal information, we analyze an asynchronous decentralized learning process in which each player independently runs a single-agent $ฮต$-greedy iterated least-squares algorithm. We prove that, despite being unable to identify the system parameters, players' learning dynamics converge almost surely to the complete-information Nash equilibrium and characterize the convergence rate. We then apply the framework to a dynamic Cournot competition with sticky prices. Numerical experiments validate the theoretical results and show that learning under limited information reduces firm profits under both low and high price stickiness, while total surplus declines and market concentration increases when price stickiness is high. Publicly revealing aggregate market output substantially accelerates convergence and mitigates these welfare losses.
Problem

Research questions and friction points this paper is trying to address.

opponent unawareness
linear-quadratic stochastic games
decentralized learning
Nash equilibrium
strategic decision-making
Innovation

Methods, ideas, or system contributions that make the work stand out.

radically uncoupled learning
linear-quadratic stochastic games
asynchronous decentralized learning
Nash equilibrium convergence
strategic obliviousness
๐Ÿ”Ž Similar Papers
No similar papers found.
D
Dantong Chu
Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong, Hong Kong, China
X
Xuefeng Gao
Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong, Hong Kong, China
Yufei Zhang
Yufei Zhang
Imperial College London
Stochastic ControlReinforcement LearningMathematical Finance