Epsilon-Nash Equilibria in History-Dependent SA-MDPs

📅 2026-09-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究了状态对抗马尔可夫决策过程中的历史依赖性问题,通过转化为等价的约束零和单边部分可观测随机博弈来计算初始状态依赖均衡的ε-近似。
📝 Abstract
We study state-adversarial Markov decision processes (SA-MDP) as a game of observation-space attacks: at each step, an agent selects an action from a received observation while an adversary$\unicode{x2014}$who knows the true state the agent is in$\unicode{x2014}$chooses a perturbed observation within a state-dependent proximity set. While existing work focuses on Markovian policies, we develop a solution concept and computational approach for SA-MDPs under history dependence. This is motivated by results showing that history dependence can materially change equilibrium outcomes and can force both the agent and the adversary to adapt their strategies. First, we prove the non-existence of universal (agnostic of the initial state distribution) history-dependent equilibrium policies. In response to this finding, our main result presents the first algorithmic route to computing $ε$-approximations of initial-state dependent equilibria. We do so by reducing SA-MDPs to a strategically equivalent constrained zero-sum one-sided partially observable stochastic game. We conclude by testing our algorithm on small analytically verifiable games and showing it scales to larger, more realistic benchmarks, including Atari Freeway rollouts with a 12-period ahead horizon.
Problem

Research questions and friction points this paper is trying to address.

SA-MDP
history-dependent
observation-space attacks
equilibrium outcomes
initial-state dependent equilibria
Innovation

Methods, ideas, or system contributions that make the work stand out.

history-dependent
epsilon-Nash equilibria
SA-MDPs
constrained zero-sum one-sided partially observable stochastic game
🔎 Similar Papers
No similar papers found.
B
Brandon Gary Kaplowitz
Department of Engineering Science, University of Oxford
D
Dominik Bohnet Zurcher
Department of Computer Science, University of Oxford
Akash Agrawal
Akash Agrawal
ML Alignment and Theory Scholars
T
Tala Jafari
Department of Computer Science, University of Oxford
Christian Schroeder de Witt
Christian Schroeder de Witt
University of Oxford
Multi-agent LearningSecuritySafety
P
Paul W. Goldberg
Department of Computer Science, University of Oxford