Continuous-time q-Learning for Jump-Diffusion Models under Tsallis Entropy

📅 2024-07-04
🏛️ arXiv.org
📈 Citations: 3
Influential: 0
📄 PDF
🤖 AI Summary
This paper investigates continuous-time reinforcement learning under jump-diffusion dynamics. Methodologically, it introduces Tsallis entropy regularization into the Q-learning framework for the first time—overcoming the inherent limitations of Gibbs-type policies arising from Shannon entropy—and establishes a martingale characterization of the Q-function via martingale representation theory and stochastic optimal control, incorporating Lagrange multipliers to derive non-Gibbsian optimal policies with compact support. Two algorithmic variants—explicit and implicit Lagrange multiplier-based continuous-time Q-learning—are proposed, enabling Actor-Critic-style alternating updates. Analytical solutions are obtained for two financial problems: optimal portfolio liquidation and nonlinear quadratic control. Numerical experiments demonstrate high stability and rapid convergence. The core contribution is the first rigorous Tsallis entropy-driven continuous-time Q-learning theoretical framework, validated for its superior modeling capability and computational efficacy under jump-diffusion dynamics.

Technology Category

Application Category

📝 Abstract
This paper studies the continuous-time reinforcement learning in jump-diffusion models by featuring the q-learning (the continuous-time counterpart of Q-learning) under Tsallis entropy regularization. Contrary to the Shannon entropy, the general form of Tsallis entropy renders the optimal policy not necessary a Gibbs measure, where the Lagrange and KKT multipliers naturally arise from some constraints to ensure the learnt policy to be a probability density function. As a consequence, the characterization of the optimal policy using the q-function also involves a Lagrange multiplier. In response, we establish the martingale characterization of the q-function under Tsallis entropy and devise two q-learning algorithms depending on whether the Lagrange multiplier can be derived explicitly or not. In the latter case, we need to consider different parameterizations of the optimal q-function and the optimal policy and update them alternatively in an Actor-Critic manner. We also study two financial applications, namely, an optimal portfolio liquidation problem and a non-LQ control problem. It is interesting to see therein that the optimal policies under the Tsallis entropy regularization can be characterized explicitly, which are distributions concentrated on some compact support. The satisfactory performance of our q-learning algorithms is illustrated in each example.
Problem

Research questions and friction points this paper is trying to address.

Study continuous-time q-learning in jump-diffusion models
Analyze Tsallis entropy regularization's impact on optimal policy
Develop q-learning algorithms for explicit and implicit Lagrange multipliers
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continuous-time q-learning with Tsallis entropy
Lagrange multiplier ensures policy probability density
Actor-Critic updates for optimal q-function
Lijun Bo
Lijun Bo
Professor, School of Mathematics and Statistics, Xidian University
Stochastic Differential EquationsMathematical Finance
Y
Yijie Huang
School of Mathematical Sciences, University of Science and Technology of China, Hefei, 230026, China
X
Xiang Yu
Department of Applied Mathematics, The Hong Kong Polytechnic University, Kowloon, Hong Kong
T
Tingting Zhang
Soochow University, Suzhou, 215006, China