$\text{DT}^2$: Decision-Targeted Digital Twins

📅 2026-06-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical limitation of conventional digital twins: when model capacity is constrained, training that minimizes single-step transition errors often yields suboptimal policy rankings, undermining effective decision-making. To overcome this, the authors propose DT², a decision-oriented digital twin training paradigm that, for the first time, explicitly optimizes for policy ranking fidelity. DT² introduces an architecture-agnostic ranking-preserving loss function that fits Q-value estimates to evaluate candidate policies and trains the twin model end-to-end to generate simulated trajectories that accurately preserve pairwise policy orderings. Experimental results demonstrate that DT² significantly improves policy ranking accuracy and reduces decision regret across diverse environments and model architectures, while maintaining high fidelity in state trajectory simulation.
📝 Abstract
A digital twin (DT) is a virtual model of a real-world system that can assist decision-making by simulating scenarios induced by different policies. However, typical machine learning-based DTs do not optimise for this use case. We prove that, when model capacity is limited, training DTs to minimise one-step transition errors can produce suboptimal models for ranking sets of policies according to a reward function. We further show that this holds empirically, even with expressive model classes. To address this, we introduce $\text{DT}^2$, a decision-targeted DT training paradigm. Firstly, $\text{DT}^2$ uses fitted Q-evaluation to estimate values of candidate policies from offline data. A DT is then trained to generate rollouts that preserve pairwise policy rankings derived from these proxy ground-truth values with an architecture-agnostic loss function. We empirically demonstrate the efficacy of our method across a range of settings and architectures. $\text{DT}^2$ consistently improves policy ranking and reduces decision regret during policy selection relative to conventional DT training, both for policies used during training and for unseen policies, while maintaining a good level of raw simulation fidelity.
Problem

Research questions and friction points this paper is trying to address.

digital twins
decision-making
policy ranking
offline reinforcement learning
simulation fidelity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Digital Twin
Decision-Targeted Learning
Policy Ranking
Fitted Q-Evaluation
Offline Reinforcement Learning
💼 Related Jobs
No related jobs found.
H
Harry Amad
Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Cambridge, United Kingdom
Mihaela van der Schaar
Mihaela van der Schaar
University of Cambridge, The Alan Turing Institute
machine learningML for healthcarecompression and streamingmulti-user networkinggame-theory