Data-Driven Merton's Strategies via Policy Randomization

📅 2023-12-19
📈 Citations: 10
Influential: 1
📄 PDF
🤖 AI Summary
This paper addresses the Merton expected utility maximization problem in an incomplete market with fully unknown dynamics: the market comprises a stock and latent state-factor processes, while investors observe only prices and instantaneous volatility—rendering factor dynamics and market parameters unidentifiable. We propose a novel continuous-time reinforcement learning (RL) framework that, for the first time, employs Gaussian policy randomization as an analytical tool—not merely for exploration—and rigorously prove that the randomized policy’s mean coincides with the optimal control of the original problem, thereby bridging RL and classical portfolio theory. We design online and offline actor-critic algorithms integrating policy improvement theorems with randomized policy analysis. Under stochastic volatility, our method substantially outperforms conventional parametric interpolation approaches. Both simulation studies and empirical tests confirm its robustness and capacity to generate alpha.
📝 Abstract
We study Merton's expected utility maximization problem in an incomplete market, characterized by a factor process in addition to the stock price process, where all the model primitives are unknown. The agent under consideration is a price taker who has access only to the stock and factor value processes and the instantaneous volatility. We propose an auxiliary problem in which the agent can invoke policy randomization according to a specific class of Gaussian distributions, and prove that the mean of its optimal Gaussian policy solves the original Merton problem. With randomized policies, we are in the realm of continuous-time reinforcement learning (RL) recently developed in Wang et al. (2020) and Jia and Zhou (2022a, 2022b, 2023), enabling us to solve the auxiliary problem in a data-driven way without having to estimate the model primitives. Specifically, we establish a policy improvement theorem based on which we design both online and offline actor-critic RL algorithms for learning Merton's strategies. A key insight from this study is that RL in general and policy randomization in particular are useful beyond the purpose for exploration -- they can be employed as a technical tool to solve a problem that cannot be otherwise solved by mere deterministic policies. At last, we carry out both simulation and empirical studies in a stochastic volatility environment to demonstrate the decisive outperformance of the devised RL algorithms in comparison to the conventional model-based, plug-in method.
Problem

Research questions and friction points this paper is trying to address.

Solves Merton's utility maximization in incomplete markets with unknown models
Uses policy randomization and RL to avoid estimating model primitives
Demonstrates RL outperforms model-based methods in stochastic volatility settings
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses policy randomization with Gaussian distributions
Employs continuous-time reinforcement learning algorithms
Solves Merton problem without estimating model primitives
🔎 Similar Papers
No similar papers found.
M
Min Dai
Department of Applied Mathematics and School of Accounting and Finance, The Hong Kong Polytechnic University, Hung Hom, Hong Kong, Kowloon
Y
Yuchao Dong
Key Laboratory of Intelligent Computing and Applications (Tongji University), Ministry of Education and School of Mathematical Sciences, Tongji University, Shanghai, China, Shanghai 200092
Y
Yanwei Jia
Department of Systems Engineering and Engineering Management, The Chinese University of Hong Kong, Shatin, Hong Kong, New Territories
X
Xun Yu Zhou
Department of Industrial Engineering and Operations Research & Data Science Institute, Columbia University, New York, USA, NY 10027