Institution profile

Amgen Inc

Industry researchnorthamerica · us
Official website
Research library10linked papers
Opportunities0open roles
Selected work

Representative Papers

Improving Power in Randomized Controlled Trials with Time-to-Event Endpoints: A Risk-Free Approach

May 26, 2026

This study addresses the challenge of safely incorporating external high-dimensional prognostic information to enhance the statistical power of randomized controlled trials with time-to-event endpoints, without introducing bias or inflating Type I error. The authors propose a two-stage framework: first constructing a prognostic score using martingale residuals and supervised learning, then integrating this score as a covariate in a nonparametric covariate-adjusted log-rank test and marginal hazard ratio estimation. This approach enables robust utilization of external prognostic data, ensuring unbiased estimation of the marginal hazard ratio and valid Type I error control even under prognostic model misspecification or population heterogeneity. Theoretical analysis shows that the variance reduction is approximately equal to the squared correlation between the prognostic score and the martingale pseudo-outcome, and simulations confirm substantial gains in efficiency—reducing required event counts and increasing power when informative prognostic signals are present.

0 citationsRead paper

Modeling Heterogeneous Mediation Effects in Survival Analysis via an Interpretable M-Learner Framework

Mar 13, 2026

This study addresses the challenge of accurately estimating heterogeneous mediation effects and identifying interpretable patient subgroups with distinct mediation pathways in clinical trials with censored outcomes. The authors propose the M-survival learner framework, which integrates causal inference, survival analysis, and machine learning to handle high-dimensional covariates and censored data. The method introduces a survival-data–oriented criterion for detecting heterogeneous mediation effects and establishes a theoretically grounded mechanism for interpretable subgroup identification, thereby supporting biomarker-informed accelerated drug approval decisions. Empirical evaluations on both a Phase III HIV clinical trial dataset and simulation studies demonstrate its strong finite-sample performance, offering reliable evidence for regulatory decision-making.

0 citationsRead paper

Controllable Generative Sandbox for Causal Inference

Mar 03, 2026

Existing causal inference simulators struggle to simultaneously preserve distributional fidelity and enable controllable causal mechanisms in mixed-type tabular data. To address this challenge, this work proposes CausalMix—a generative framework based on variational autoencoders that models complex data distributions through a Gaussian mixture latent prior and data-type-specific decoders, while explicitly parameterizing key causal factors such as overlap, unobserved confounding strength, and treatment effect heterogeneity. CausalMix is the first method to unify realistic mixed-data generation with factorized, interpretable control over causal mechanisms, allowing independent adjustment of each causal component during simulation design. Empirical evaluations demonstrate that CausalMix achieves state-of-the-art distributional fidelity across multiple benchmarks while providing stable and fine-grained causal control, and it has been successfully applied to estimator evaluation and power analysis in a comparative study of prostate cancer treatments.

0 citationsRead paper

Unsupervised dense random survival forests identify interpretable patient profiles with heterogeneous treatment benefit

Jan 04, 2026

This study addresses the challenge of precision oncology by identifying patient subgroups with heterogeneous treatment effects in randomized clinical trials of experimental cancer therapies. To this end, we propose an unsupervised machine learning approach that constructs ultra-dense random survival forests—comprising up to 100,000 trees—and introduces the first unsupervised splitting criterion explicitly designed to model treatment–covariate interactions for heterogeneity of treatment effect. The method maintains a low false positive rate (Type I error <1%) while offering high interpretability and robustness. Experiments on both simulated data and real-world Phase III clinical trials demonstrate its ability to accurately distinguish scenarios with and without treatment effect heterogeneity and to effectively identify patient subgroups exhibiting significantly different therapeutic responses.

0 citationsRead paper

Integrating RCTs, RWD, AI/ML and Statistics: Next-Generation Evidence Synthesis

Nov 24, 2025

This study addresses key challenges in evidence generation: the high cost and limited generalizability of randomized controlled trials (RCTs); the low causal validity of real-world data (RWD); and the insufficient interpretability and statistical rigor of AI/ML models. We propose a novel evidence synthesis framework integrating causal inference, privacy-preserving machine learning, uncertainty quantification, and small-sample statistics. Methodologically, we develop a “causal reasoning roadmap” enabling RCT result extrapolation, AI-embedded analysis, hybrid trial design, and dynamic integration of short-term RCTs with long-term RWD. Our contribution lies in deeply embedding interpretable AI within the causal modeling pipeline—thereby ensuring statistical robustness, regulatory acceptability, and policy relevance simultaneously. The framework significantly enhances external validity, transparency, and regulatory applicability of evidence, establishing a reproducible, verifiable next-generation paradigm for regulatory science.

0 citationsRead paper
Recent publications

Latest Papers

Improving Power in Randomized Controlled Trials with Time-to-Event Endpoints: A Risk-Free Approach

May 26, 2026

This study addresses the challenge of safely incorporating external high-dimensional prognostic information to enhance the statistical power of randomized controlled trials with time-to-event endpoints, without introducing bias or inflating Type I error. The authors propose a two-stage framework: first constructing a prognostic score using martingale residuals and supervised learning, then integrating this score as a covariate in a nonparametric covariate-adjusted log-rank test and marginal hazard ratio estimation. This approach enables robust utilization of external prognostic data, ensuring unbiased estimation of the marginal hazard ratio and valid Type I error control even under prognostic model misspecification or population heterogeneity. Theoretical analysis shows that the variance reduction is approximately equal to the squared correlation between the prognostic score and the martingale pseudo-outcome, and simulations confirm substantial gains in efficiency—reducing required event counts and increasing power when informative prognostic signals are present.

0 citationsRead paper

Modeling Heterogeneous Mediation Effects in Survival Analysis via an Interpretable M-Learner Framework

Mar 13, 2026

This study addresses the challenge of accurately estimating heterogeneous mediation effects and identifying interpretable patient subgroups with distinct mediation pathways in clinical trials with censored outcomes. The authors propose the M-survival learner framework, which integrates causal inference, survival analysis, and machine learning to handle high-dimensional covariates and censored data. The method introduces a survival-data–oriented criterion for detecting heterogeneous mediation effects and establishes a theoretically grounded mechanism for interpretable subgroup identification, thereby supporting biomarker-informed accelerated drug approval decisions. Empirical evaluations on both a Phase III HIV clinical trial dataset and simulation studies demonstrate its strong finite-sample performance, offering reliable evidence for regulatory decision-making.

0 citationsRead paper

Controllable Generative Sandbox for Causal Inference

Mar 03, 2026

Existing causal inference simulators struggle to simultaneously preserve distributional fidelity and enable controllable causal mechanisms in mixed-type tabular data. To address this challenge, this work proposes CausalMix—a generative framework based on variational autoencoders that models complex data distributions through a Gaussian mixture latent prior and data-type-specific decoders, while explicitly parameterizing key causal factors such as overlap, unobserved confounding strength, and treatment effect heterogeneity. CausalMix is the first method to unify realistic mixed-data generation with factorized, interpretable control over causal mechanisms, allowing independent adjustment of each causal component during simulation design. Empirical evaluations demonstrate that CausalMix achieves state-of-the-art distributional fidelity across multiple benchmarks while providing stable and fine-grained causal control, and it has been successfully applied to estimator evaluation and power analysis in a comparative study of prostate cancer treatments.

0 citationsRead paper

Unsupervised dense random survival forests identify interpretable patient profiles with heterogeneous treatment benefit

Jan 04, 2026

This study addresses the challenge of precision oncology by identifying patient subgroups with heterogeneous treatment effects in randomized clinical trials of experimental cancer therapies. To this end, we propose an unsupervised machine learning approach that constructs ultra-dense random survival forests—comprising up to 100,000 trees—and introduces the first unsupervised splitting criterion explicitly designed to model treatment–covariate interactions for heterogeneity of treatment effect. The method maintains a low false positive rate (Type I error <1%) while offering high interpretability and robustness. Experiments on both simulated data and real-world Phase III clinical trials demonstrate its ability to accurately distinguish scenarios with and without treatment effect heterogeneity and to effectively identify patient subgroups exhibiting significantly different therapeutic responses.

0 citationsRead paper

Integrating RCTs, RWD, AI/ML and Statistics: Next-Generation Evidence Synthesis

Nov 24, 2025

This study addresses key challenges in evidence generation: the high cost and limited generalizability of randomized controlled trials (RCTs); the low causal validity of real-world data (RWD); and the insufficient interpretability and statistical rigor of AI/ML models. We propose a novel evidence synthesis framework integrating causal inference, privacy-preserving machine learning, uncertainty quantification, and small-sample statistics. Methodologically, we develop a “causal reasoning roadmap” enabling RCT result extrapolation, AI-embedded analysis, hybrid trial design, and dynamic integration of short-term RCTs with long-term RWD. Our contribution lies in deeply embedding interpretable AI within the causal modeling pipeline—thereby ensuring statistical robustness, regulatory acceptability, and policy relevance simultaneously. The framework significantly enhances external validity, transparency, and regulatory applicability of evidence, establishing a reproducible, verifiable next-generation paradigm for regulatory science.

0 citationsRead paper