Control-Oriented Scenario Tree Construction through Reinforcement Learning

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Traditional scenario tree construction methods prioritize probabilistic fidelity but often fail to guarantee strong downstream control performance. This work proposes a novel control-performance-driven paradigm, formulating scenario allocation as an attention-based policy optimization problem. By leveraging reinforcement learning, the approach directly maximizes closed-loop control returns and generates compact branches that emphasize high-impact events under a fixed tree structure. The method incorporates an asymmetric critic to stabilize training and is evaluated against baselines such as Wasserstein reduction. In a risk-averse battery arbitrage task, it consistently achieves the highest returns across varying forecast ensemble sizes, significantly outperforming classical scenario reduction and certainty-equivalent control while demonstrating superior tail-risk management.
📝 Abstract
Multistage stochastic model predictive control (MPC) handles uncertainty by optimizing over a scenario tree, a finite branching approximation of future outcomes constructed from sampled forecasts. To build such a tree, conventional methods focus on matching the underlying probability distribution---e.g., via Wasserstein-based scenario reduction---but improved distributional accuracy does not necessarily yield better control performance. We propose a control-oriented approach that learns scenario tree construction directly from its impact on downstream decisions. Fixing the tree topology, we formulate tree construction as a sequential assignment of sampled scenarios to leaves. This assignment is parameterized by an attention-based policy over the scenario set and trained using reinforcement learning, with closed-loop control profit as the objective. Training is stabilized by an asymmetric critic that leverages realized future trajectories. We evaluate the method on a risk-averse battery arbitrage problem. Across a range of forecast set sizes, the learned construction consistently achieves the highest profit, outperforming classical forward and backward reduction methods and certainty-equivalent (single-trajectory forecast) control. The learned policy also exhibits greater robustness on challenging instances, consistently demonstrating better tail-risk characteristics. Analysis of the resulting trees indicates that our method constructs compact, selectively branching structures that capture high-impact events while keeping most trajectories nearly deterministic. These findings highlight that the value of a scenario tree depends critically on the decisions it supports, and provide an effective framework to train scenario tree constructors merely based on the closed-loop control optimization signal.
Problem

Research questions and friction points this paper is trying to address.

scenario tree
stochastic MPC
control performance
uncertainty handling
decision-oriented construction
Innovation

Methods, ideas, or system contributions that make the work stand out.

control-oriented scenario tree
reinforcement learning
multistage stochastic MPC
attention-based policy
closed-loop control optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Fabio Pavirani
IDLab Ghent university – imec, Technologiepark Zwijnaarde 126, 9052 Gent, Belgium
B
Bert Claessens
Beebop.ai, Belgium
Pierre Pinson
Pierre Pinson
Imperial College London
ForecastingGame theoryDecision-making under uncertainty
Chris Develder
Chris Develder
Ghent University - imec
Smart GridInformation ExtractionOptical Networks