Online Policy Evaluation for MDPs with Dynamic UBSR Measures
This work addresses the limitations of existing risk-aware reinforcement learning methods for policy evaluation, which are often confined to specific risk measures or rely on simulators, thereby hindering their applicability in fully online settings. Focusing on Markov decision processes under the dynamic utility-based shortfall risk (UBSR) measure, the study introduces UBSR-TD—an efficient online policy evaluation algorithm based on linear function approximation—by extending the risk-neutral temporal difference (TD) algorithm through a tailored loss function, along with an accelerated variant. Theoretical analysis establishes its almost sure convergence, while numerical experiments confirm its empirical effectiveness. Furthermore, the approach demonstrates practical utility in managing perishable inventory with uncertain shelf life, thereby overcoming key applicability barriers in online risk-aware policy evaluation.