Threshold Structure of Optimal Policies in Restart POMDPs

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses optimal control in restart-type partially observable Markov decision processes (Restart POMDPs), where the system state is observable only at restart epochs. By introducing a sufficient statistic—comprising the last observed state and the time elapsed since the most recent restart—the problem is transformed into a fully observable MDP. Under both discounted and average cost criteria, the authors establish that when the one-stage cost satisfies a deterioration condition and the transition kernel exhibits monotonicity, the optimal policy admits a time-threshold structure, with thresholds non-increasing in the observed state. The analysis leverages dimensionality reduction via sufficient statistics, stochastic monotonicity, geometric ergodicity, vanishing discount techniques, and uniform boundedness of the relative value function. These results provide a theoretical foundation and a structured policy representation for efficiently solving Restart POMDPs.
📝 Abstract
We study a Restart POMDP (Partially Observable Markov Decision Process) on a general Borel state space, where the controller either lets the hidden state evolve unobserved or restarts the system and observes the new state. Exploiting a sufficient-statistic representation consisting of the last observed state and the elapsed time since restart, we reduce the problem to a fully observed MDP. Under a natural one-step cost deterioration condition, we prove that optimal policies have a threshold structure in the elapsed time for both the discounted and total undiscounted cost criteria. When the state space is partially ordered and the kernel is stochastically monotone, we further show that the optimal threshold is nonincreasing in the state. For the average cost criterion, under additional assumptions of geometric ergodicity and domination of the transient gain, we establish analogous threshold results via the vanishing discount approach, after showing the uniform boundedness of the optimal thresholds and relative value functions.
Problem

Research questions and friction points this paper is trying to address.

POMDP
threshold policy
restart
optimal control
Markov decision process
Innovation

Methods, ideas, or system contributions that make the work stand out.

Threshold Policy
Restart POMDP
Sufficient Statistic
Stochastic Monotonicity
Vanishing Discount
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Konstantin Avrachenkov
Konstantin Avrachenkov
Director of Research, INRIA Sophia Antipolis
Applied ProbabilityMarkov ChainsSingular PerturbationsNetworksMachine Learning
A
Alexey Piunovskiy
Department of Mathematical Sciences, University of Liverpool, UK
Y
Yi Zhang
School of Mathematics, University of Birmingham, UK