Policy Optimization and Statistical Inference for Online Contextual Matrix Games

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了在线决策中动态环境和策略互动的问题,提出OnGameLearn算法,结合上下文信息进行多玩家在线游戏的学习,平衡探索与利用。
📝 Abstract
Online decision making often requires navigating a landscape shaped by both dynamic contexts and strategic interactions. In competitive pricing, for example, hotels must account for both dynamic contextual factors and rivals' strategic responses. Existing approaches address only part of this challenge: contextual bandits optimize single-agent decisions using observable features but ignore multi-player interactions, while online matrix games capture strategic behavior through Nash equilibrium but assume fixed payoffs, ignoring contextual information. How should agents act then when strategic payoffs evolve with contextual signals? We introduce \emph{online contextual matrix games} to integrate contextual information into multi-player online games. We further propose \emph{OnGameLearn}, an online learning algorithm that efficiently balances exploration and exploitation across both player actions and contexts. This approach comes with statistical guarantees: tail bounds for the estimated payoff matrix, the convergence of the estimated Nash equilibrium, the asymptotic normality of the parameter estimators, and the sublinear regret bound. We also develop the notion of \emph{policy value} in matrix games and develop a doubly robust, $\sqrt{T}$-consistent estimator for it. Across simulated studies and a real-world hotel pricing application, we find that OnGameLearn effectively navigates the intertwined challenges of strategic and contextual decision-making.
Problem

Research questions and friction points this paper is trying to address.

online decision making
contextual information
strategic interactions
Nash equilibrium
dynamic payoffs
Innovation

Methods, ideas, or system contributions that make the work stand out.

online contextual matrix games
OnGameLearn
policy value
statistical guarantees
🔎 Similar Papers
No similar papers found.
L
Liner Xiang
Department of Statistics, University of California, Irvine
Yixin Wang
Yixin Wang
University of Michigan
Bayesian statisticsMachine Learning
H
Hengrui Cai
Department of Statistics, University of California, Irvine