Reinforcement Learning for Continuous-Time Jump Markov Decision Processes with Applications to Network Dynamic Pricing

📅 2026-08-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文研究了连续时间跳跃马尔可夫决策过程中的强化学习问题,通过建立熵正则化控制问题和开发无模型q学习算法来解决具有通用离散状态空间的应用场景中的探索与利用平衡问题。
📝 Abstract
We study reinforcement learning (RL) in Continuous-Time Jump Markov Decision Processes (CTJMDPs) featuring general discrete state spaces (which need not possess a vector space structure) and continuous/discrete action spaces. The setup covers many well-known applications in operations such as multi-product dynamic pricing with capacitated resources (Gallego and van Ryzin 1997). To model the exploration-exploitation tradeoff, we formulate an entropy-regularized continuous-time control problem with stochastic policies. Recent continuous-time RL techniques such as $q$-learning for controlled diffusions in (Jia and Zhou 2023) focus on continuous state spaces $\mathbb{R}^d$ and rely heavily on semimartingale theory in $\mathbb{R}^d$ for their theoretical analysis. Consequently, their methods cannot be directly applied to CTJMDPs with general discrete state spaces, which may lack the algebraic addition and subtraction structures inherent to Euclidean spaces. To bridge this gap, we establish the theoretical foundations of $q$-learning for CTJMDPs and develop model-free $q$-learning algorithms. Compared to naïve time discretization and approximating CTJMDPs using discrete-time MDPs, our approach has several conceptual and empirical benefits. Numerical experiments in network dynamic pricing (Gallego and van Ryzin 1997) show that our proposed RL algorithm reliably learns near-optimal policies and consistently outperforms standard benchmark methods, demonstrating superior solution quality and effective scalability to large-scale network instances.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Continuous-Time Jump Markov Decision Processes
Network Dynamic Pricing
Exploration-Exploitation Tradeoff
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Continuous-Time Jump Markov Decision Processes
Entropy-Regularization
Model-Free Q-Learning
Network Dynamic Pricing
H
Huiling Meng
Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong, Hong Kong, China
Ningyuan Chen
Ningyuan Chen
Department of Management, UTM & Rotman School of Management, University of Toronto
Revenue ManagementOnline LearningOperations ManagementBusiness Analytics
X
Xuefeng Gao
Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong, Hong Kong, China