🤖 AI Summary
This study addresses the challenge of precision oncology by identifying patient subgroups with heterogeneous treatment effects in randomized clinical trials of experimental cancer therapies. To this end, we propose an unsupervised machine learning approach that constructs ultra-dense random survival forests—comprising up to 100,000 trees—and introduces the first unsupervised splitting criterion explicitly designed to model treatment–covariate interactions for heterogeneity of treatment effect. The method maintains a low false positive rate (Type I error <1%) while offering high interpretability and robustness. Experiments on both simulated data and real-world Phase III clinical trials demonstrate its ability to accurately distinguish scenarios with and without treatment effect heterogeneity and to effectively identify patient subgroups exhibiting significantly different therapeutic responses.
📝 Abstract
Precision oncology aims to prescribe the optimal cancer treatment to the right patients, maximizing therapeutic benefits. However, identifying patient subgroups that may benefit more from experimental cancer treatments based on randomized clinical trials presents a significant analytical challenge. To address this, we introduce a novel unsupervised machine learning approach based on very dense random survival forests (up to 100,000 trees), equipped with a new splitting rule that explicitly targets treatment-effect heterogeneity. This method is robust, interpretable, and effectively identifies responsive subgroups. Extensive simulations confirm its ability to detect heterogeneous patient responses and distinguish between datasets with and without heterogeneity, while maintaining a stringent Type I error rate of 1%. We further validate its performance using Phase III randomized clinical trial datasets, demonstrating significant patient heterogeneity in treatment response based on baseline characteristics.