Adaptive Policy Portfolios for Robust Markov Decision Processes

📅 2026-08-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过离线合成有限的记忆随机策略集和轻量级在线选择器,以适应部分可识别的固定未知动态,优化鲁棒马尔可夫决策过程。
📝 Abstract
Robust Markov decision processes optimize one policy against a set of plausible transition functions. This can be conservative when the unknown dynamics are fixed and become partially identifiable after deployment. We study adaptive policy portfolios: finite sets of memoryless randomized policies synthesized offline and paired with a lightweight online selector. Robust regret is a natural measure of portfolio quality: for each plausible environment, it measures the loss of the best portfolio member relative to the policy that would have been optimal had that environment been known. Related regret objectives were studied by Ghavamzadeh et al. (2016) with an emphasis on approximations and relaxations for safe policy improvement. We give a complexity-theoretic account of portfolio certification and synthesis. Certifying a given portfolio is $\forall\mathbb{R}$-complete already for deterministic portfolios in acyclic (s,a)-rectangular RMDPs. Synthesizing a portfolio of unary-bounded size is $\exists\forall\mathbb{R}$-complete for general rational polytopes, even with fixed discount and acyclic dynamics. The single-policy case is already hard, both combinatorially and algebraically. Finally, we present an offline portfolio construction that is amenable to runtime specialization.
Problem

Research questions and friction points this paper is trying to address.

Adaptive Policy Portfolios
Robust Markov Decision Processes
Robust Regret
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Policy Portfolios
Robust Markov Decision Processes
Regret Minimization
Complexity Analysis
Offline Construction
🔎 Similar Papers
No similar papers found.