🤖 AI Summary
This study addresses the challenge of reliably estimating treatment effects in single-arm clinical trials due to the absence of a concurrent control group. The authors propose a machine learning–based digital twin approach that constructs a synthetic control arm by generating personalized predictions of disease progression for untreated patients. Integrating doubly robust estimation with the U.S. FDA’s artificial intelligence/ML guidance principles, the method leverages historical data for model development and facilitates sample size calculations. Reanalysis of real-world trial data in amyotrophic lateral sclerosis and Huntington’s disease demonstrates that the proposed framework substantially improves the accuracy and robustness of treatment effect estimation, offering a flexible and scalable paradigm for causal inference in single-arm trials.
📝 Abstract
Single-arm trials are an important study design for evaluating drug efficacy and safety without enrolling patients into a control arm. Although they do not provide the gold-standard evidence of randomized controlled trials, they are increasingly used in clinical development as they offer an efficient, ethical, and practical alternative. A wide variety of approaches can be used to construct control comparators and estimate treatment effects, from fixed comparators informed by clinical knowledge to data-based and model-based patient-level comparators, also known as synthetic controls. Powerful and flexible machine learning models can allow outcome-model-based synthetic controls to overcome key limitations of direct data-based approaches, yield more robust estimates of treatment effects, and provide a principled way to incorporate corrections or encode additional assumptions when external data are not directly comparable. In this work, we argue that outcome-model-based synthetic control arms are an important tool for single-arm trials. We focus on digital twins, personalized predictions of disease progression generated from machine learning models trained on historical datasets, which naturally leverage these flexible approaches. We review doubly robust estimators, present power and sample size formulas, and discuss trade-offs in selecting historical data for training and analysis. We also outline practical considerations for deploying digital twins within the framework of recent FDA draft guidance on the use of artificial intelligence in drug development. Finally, we reanalyze data from trials in amyotrophic lateral sclerosis and Huntington's disease to demonstrate the proposed methods.