🤖 AI Summary
This work addresses the computational challenges in non-Markovian stochastic optimal control, where the value function is governed by a stochastic Hamilton–Jacobi–Bellman (SHJB) equation whose measurability-induced randomness impedes tractable solution. Under the assumption of control-independent stochastic integral coefficients, the paper introduces the first policy iteration framework tailored to semilinear SHJB equations. The approach iteratively linearizes the original problem into a sequence of linear equations, which are then solved using tools from stochastic analysis and functional approximation. Theoretical analysis establishes that the resulting sequence of approximations converges monotonically in the mean-square sense and exhibits exponential convergence rates, thereby significantly enhancing both the feasibility and efficiency of computing value functions in non-Markovian settings.
📝 Abstract
This paper is concerned with the non-Markovian stochastic optimal control problems in which the value function is a random field characterized by a stochastic Hamilton-Jacobi-Bellman (SHJB) equation. When the stochastic integration coefficients are not controlled, the SHJB equation takes a semilinear form, which is subject to computational challenges compared to the Markovian case due to the measurable randomness. We introduce a policy-iteration algorithm based on successive linearization that reduces the nonlinear SHJB equation to a sequence of linear ones. Furthermore, we prove that the resulting approximation sequence converges monotonically to the value function in the mean-square sense with an exponential rate.