Learning to Control Coupled-Dynamics Environments with Joint Markov Decision Processes

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文解决了在耦合动力学环境中使用联合马尔可夫决策过程进行最优控制的问题,通过定义非参数分布贝尔曼最优算子,并证明其收敛性。
📝 Abstract
Coupled-dynamics environments expose the one-step outcomes that would follow from several possible counterfactual actions under a common realization of exogenous randomness. The ordinary Markov decision process formalism allows one to reason about the marginal law of each action but discards dependence across these counterfactual outcomes. The Joint Markov decision process (JMDP) formalism preserves that dependence. Prior work established the formalism and solved the fixed-policy joint moment evaluation problem in JMDPs. This paper develops optimal-control methods. We define a nonparametric distributional Bellman optimality operator for JMDPs, and prove that when the induced marginal MDP has a unique optimal policy, its iterates converge in Wasserstein distance to the optimal joint return law. For the first two moments, we establish convergence under a weaker condition that permits several mean-optimal actions as long as their tie resolutions share a second-moment fixed point. We also derive sampled targets for neural approximation.
Problem

Research questions and friction points this paper is trying to address.

Coupled-dynamics
Markov decision process
Optimal control
Joint Markov decision process
Innovation

Methods, ideas, or system contributions that make the work stand out.

Joint Markov Decision Process (JMDP)
Nonparametric Distributional Bellman Optimality Operator
Wasserstein Distance
Convergence
Neural Approximation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.