Scalable Pontryagin-Guided Adjoint-to-Control Recovery for Constrained Dynamic Portfolio Choice

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of adjoint variable recovery and scalable control in continuous-time dynamic portfolio selection with smoothing constraints by proposing an "Adjoint-to-Control" framework. By establishing a theoretical correspondence between Open-Loop Backpropagation Through Time (OL-BPTT) and Pontryagin’s Maximum Principle (PMP), and integrating Differentiable Path Optimization (DPO) for path sensitivity, the method employs nested regression to estimate martingale inputs and solve local generalized Hamiltonian problems. This approach enables efficient and precise recovery of optimal strategies under high-dimensional constraints. Benchmarks involving one hundred assets demonstrate an adjoint error of only 0.46% and a strategy RMSE below 8.5×10⁻³, while significantly reducing KKT residuals. These results validate the method's accuracy and practicality for high-dimensional financial control applications.
📝 Abstract
We develop a scalable adjoint-to-control framework for continuous-time portfolio choice under smooth pointwise constraints. A feasible direct-policy-optimization (DPO) policy supplies rollouts; after training, fixed-latent open-loop BPTT (OL-BPTT) yields first- and second-order pathwise sensitivities, whose conditional projections produce adapted adjoint inputs. A nested antithetic common-random-number regression estimates the shifted wealth-row martingale input, and deployment solves the local generalized-Hamiltonian problem by an exact QP for quadratic-affine blocks or by a log barrier otherwise. We prove an OL-BPTT--PMP correspondence retaining orthogonal projection residuals, a local barrier--KKT approximation, and a performance-to-adjoint bridge under local quadratic growth. In an n=100 constrained Merton benchmark, the learned first adjoint has 0.46% mean relative error; under the analytical policy, the first adjoint, wealth curvature, and Brownian coefficient have nRMSEs of 0.031%, 0.035%, and 0.326%. In a terminal-only predictable-return CRRA benchmark, wealth homogeneity gives $P^{X\bullet,*}=D^2_{X\bullet}V$ and $ζ^{X,*}=0$. At the primary $512\times16$ projection budget, the full estimated-shift decoder has policy RMSE below $8.5\times10^{-3}$, while the benchmark-specific zero-shift oracle is below $2\times10^{-4}$. Across constraint, factor, barrier, and switching-region audits, recovery sharply reduces local KKT residuals and remains practical for portfolios with up to 100 risky assets.
Problem

Research questions and friction points this paper is trying to address.

Constrained Dynamic Portfolio Choice
Adjoint-to-Control Recovery
Continuous-Time Optimization
Scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adjoint-to-Control Recovery
Open-Loop BPTT
Constrained Dynamic Portfolio Choice
Antithetic Common-Random-Number Regression
Generalized-Hamiltonian Problem
🔎 Similar Papers
No similar papers found.