🤖 AI Summary
This study investigates the set of achievable value functions in infinite-horizon partially observable Markov decision processes (POMDPs) under memoryless stochastic policies. Addressing the long-standing lack of a precise mathematical characterization of this set, the work establishes for the first time that it forms a semialgebraic set. Specifically, it explicitly constructs a system of polynomial inequalities—derived from the system dynamics, observation kernel, and reward structure—that fully describes the feasible value functions, thereby revealing the nonlinear constraints and intricate geometric structure induced by partial observability. This result generalizes the classical finding that the value function set in fully observable MDPs is polyhedral, clarifies the dependence of achievable values on the initial state distribution, and uncovers novel phenomena such as isolated locally optimal policies, thus providing a new theoretical foundation for POMDP policy optimization.
📝 Abstract
We study the geometry of feasible value functions in infinite-horizon partially observable Markov decision processes (POMDPs) under memoryless stochastic policies. Our main contribution is a characterization of the feasible set of value functions as a semi-algebraic set, defined by explicit polynomial inequalities determined by the transition dynamics, observation kernel, and reward structure of the POMDP. This result extends prior work for fully observable Markov decision processes, where the feasible set is known to be a polytope, to the substantially more intricate partially observable setting. In contrast to the polyhedral structure arising in MDPs, partial observability induces fundamentally nonlinear constraints, leading to a richer and more complex geometric structure. Our geometric characterization provides new insight into the landscape of policy optimization in both MDPs and POMDPs, and reveals qualitative phenomena unique to partial observability, including the emergence of isolated local maximizers of the long-term reward and their dependence on the initial state distribution.