Institution profile

St. Jude Children's Research Hospital

Academic institutionnorthamerica · us
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

Interim Monitoring as an Information-Time Alignment Problem: The WCR Framework for Time-to-Event Trials

Jun 13, 2026

This study addresses the longstanding challenge in time-to-event trials of balancing inferential maturity with practical feasibility during interim monitoring, where conventional event-driven or enrollment-driven approaches suffer from inherent limitations. The authors propose the WCR framework, which reframes interim monitoring as an information–time alignment problem. By fixing the cohort size and calibrating follow-up requirements, WCR enables continued enrollment while synchronizing interim analyses with information maturity, reserving later-enrolled patients for the final analysis. The framework explicitly distinguishes follow-up constraints between landmark survival estimators and proportional hazards models, jointly calibrates design parameters and decision thresholds, and integrates constrained optimization, simulation-based calibration, and Bayesian methods, implemented in the open-source R package WCRBayesDesign. In simulations of rare pediatric oncology trials, WCR substantially improves the stability and interpretability of interim analysis timing while rigorously controlling Type I error and maintaining power, outperforming existing strategies.

0 citationsRead paper

A Tutorial for Evaluating Cure Model Appropriateness

May 06, 2026

In survival analysis, traditional models assume all individuals will eventually experience the event of interest. However, advances in therapeutics have led to multiple clinical contexts with potentially curative therapies, and in these contexts, certain individuals may never experience the event. Statisticians have developed cure models as a methodology to address this challenge. Nonetheless, despite significant statistical advances in cure models, we have seen more limited uptake in biomedical applications, and we hypothesize that this is caused by limited guidance in the appropriate application of cure models. Cure models require specific identifiability conditions for valid parameter estimation, and previous reports have demonstrated significant issues with the inappropriate application of cure models. Existing tutorials for cure models focus on model implementation and either assume or provide only limited guidance on whether cure modeling is appropriate for the given dataset. This tutorial addresses this gap by describing a systematic procedure that integrates clinical judgment, visual inspection of Kaplan-Meier curves, and quantitative evaluation. We provide a worked example using data from a randomized clinical trial in acute myeloid leukemia, and we also summarize findings from a series of other datasets of hematopoietic cell transplantation to suggest broad practical guidance for choosing to apply cure models. By systematically evaluating cure model appropriateness before fitting these models, researchers can achieve more reliable survival analysis and improved clinical decision-making.

0 citationsRead paper

A comparative study of two-sample hypothesis tests in the presence of long-term survivors

May 04, 2026

This study addresses the limitations of conventional two-sample time-to-event tests in the presence of long-term survivors (L-TS) under non-proportional hazards, where ignoring the cured fraction can substantially reduce statistical power and where the impact of follow-up duration remains inadequately characterized. Through extensive Monte Carlo simulations under neutral scenarios, the authors systematically compare the type I error and power of traditional tests (e.g., log-rank), non-proportional hazards adjustments, and correctly specified parametric cure models across varying sample sizes, follow-up durations, and effect magnitudes. They find that when both groups contain L-TS, the power of conventional methods varies non-monotonically with follow-up time, whereas parametric cure models exhibit monotonically increasing power. A numerical tool is proposed to predict this non-monotonic behavior, thereby informing optimal follow-up design. The results demonstrate that parametric cure models offer superior performance under prolonged follow-up, establishing a new paradigm for trial design in settings with L-TS.

0 citationsRead paper

When Privacy Isn't Synthetic: Hidden Data Leakage in Generative AI Models

Dec 05, 2025

This paper uncovers a previously overlooked membership inference risk in generative AI–synthesized data: even without explicit memorization of training samples, structural overlap between the original and synthetic data manifolds can leak individual membership information. To exploit this vulnerability, we propose a clustering-center–based black-box inference attack—leveraging unsupervised clustering and density estimation to identify dense neighborhoods in the synthetic data distribution, constructing proxy representations of training samples, and enabling high-accuracy membership inference via black-box queries. Crucially, our approach attributes privacy leakage to distributional neighborhood alignment rather than pointwise sample memorization, thereby challenging conventional privacy evaluation paradigms. Experiments across sensitive domains—including healthcare and finance—demonstrate significant membership leakage in synthetically generated data, even when trained with differential privacy, revealing fundamental limitations of existing privacy-preserving mechanisms.

0 citationsRead paper
Recent publications

Latest Papers

Interim Monitoring as an Information-Time Alignment Problem: The WCR Framework for Time-to-Event Trials

Jun 13, 2026

This study addresses the longstanding challenge in time-to-event trials of balancing inferential maturity with practical feasibility during interim monitoring, where conventional event-driven or enrollment-driven approaches suffer from inherent limitations. The authors propose the WCR framework, which reframes interim monitoring as an information–time alignment problem. By fixing the cohort size and calibrating follow-up requirements, WCR enables continued enrollment while synchronizing interim analyses with information maturity, reserving later-enrolled patients for the final analysis. The framework explicitly distinguishes follow-up constraints between landmark survival estimators and proportional hazards models, jointly calibrates design parameters and decision thresholds, and integrates constrained optimization, simulation-based calibration, and Bayesian methods, implemented in the open-source R package WCRBayesDesign. In simulations of rare pediatric oncology trials, WCR substantially improves the stability and interpretability of interim analysis timing while rigorously controlling Type I error and maintaining power, outperforming existing strategies.

0 citationsRead paper

A Tutorial for Evaluating Cure Model Appropriateness

May 06, 2026

In survival analysis, traditional models assume all individuals will eventually experience the event of interest. However, advances in therapeutics have led to multiple clinical contexts with potentially curative therapies, and in these contexts, certain individuals may never experience the event. Statisticians have developed cure models as a methodology to address this challenge. Nonetheless, despite significant statistical advances in cure models, we have seen more limited uptake in biomedical applications, and we hypothesize that this is caused by limited guidance in the appropriate application of cure models. Cure models require specific identifiability conditions for valid parameter estimation, and previous reports have demonstrated significant issues with the inappropriate application of cure models. Existing tutorials for cure models focus on model implementation and either assume or provide only limited guidance on whether cure modeling is appropriate for the given dataset. This tutorial addresses this gap by describing a systematic procedure that integrates clinical judgment, visual inspection of Kaplan-Meier curves, and quantitative evaluation. We provide a worked example using data from a randomized clinical trial in acute myeloid leukemia, and we also summarize findings from a series of other datasets of hematopoietic cell transplantation to suggest broad practical guidance for choosing to apply cure models. By systematically evaluating cure model appropriateness before fitting these models, researchers can achieve more reliable survival analysis and improved clinical decision-making.

0 citationsRead paper

A comparative study of two-sample hypothesis tests in the presence of long-term survivors

May 04, 2026

This study addresses the limitations of conventional two-sample time-to-event tests in the presence of long-term survivors (L-TS) under non-proportional hazards, where ignoring the cured fraction can substantially reduce statistical power and where the impact of follow-up duration remains inadequately characterized. Through extensive Monte Carlo simulations under neutral scenarios, the authors systematically compare the type I error and power of traditional tests (e.g., log-rank), non-proportional hazards adjustments, and correctly specified parametric cure models across varying sample sizes, follow-up durations, and effect magnitudes. They find that when both groups contain L-TS, the power of conventional methods varies non-monotonically with follow-up time, whereas parametric cure models exhibit monotonically increasing power. A numerical tool is proposed to predict this non-monotonic behavior, thereby informing optimal follow-up design. The results demonstrate that parametric cure models offer superior performance under prolonged follow-up, establishing a new paradigm for trial design in settings with L-TS.

0 citationsRead paper

When Privacy Isn't Synthetic: Hidden Data Leakage in Generative AI Models

Dec 05, 2025

This paper uncovers a previously overlooked membership inference risk in generative AI–synthesized data: even without explicit memorization of training samples, structural overlap between the original and synthetic data manifolds can leak individual membership information. To exploit this vulnerability, we propose a clustering-center–based black-box inference attack—leveraging unsupervised clustering and density estimation to identify dense neighborhoods in the synthetic data distribution, constructing proxy representations of training samples, and enabling high-accuracy membership inference via black-box queries. Crucially, our approach attributes privacy leakage to distributional neighborhood alignment rather than pointwise sample memorization, thereby challenging conventional privacy evaluation paradigms. Experiments across sensitive domains—including healthcare and finance—demonstrate significant membership leakage in synthetically generated data, even when trained with differential privacy, revealing fundamental limitations of existing privacy-preserving mechanisms.

0 citationsRead paper