🤖 AI Summary
This study addresses the limitations of conventional two-sample time-to-event tests in the presence of long-term survivors (L-TS) under non-proportional hazards, where ignoring the cured fraction can substantially reduce statistical power and where the impact of follow-up duration remains inadequately characterized. Through extensive Monte Carlo simulations under neutral scenarios, the authors systematically compare the type I error and power of traditional tests (e.g., log-rank), non-proportional hazards adjustments, and correctly specified parametric cure models across varying sample sizes, follow-up durations, and effect magnitudes. They find that when both groups contain L-TS, the power of conventional methods varies non-monotonically with follow-up time, whereas parametric cure models exhibit monotonically increasing power. A numerical tool is proposed to predict this non-monotonic behavior, thereby informing optimal follow-up design. The results demonstrate that parametric cure models offer superior performance under prolonged follow-up, establishing a new paradigm for trial design in settings with L-TS.
📝 Abstract
Time-to-event data with long-term survivors (L-TS), subjects who never experience the event, have been reported in multiple areas of oncology as therapies have improved. Conventional two-sample tests ignore L-TS, but alternatives have been developed in the cure models literature. Because L-TS can induce non-proportional hazards (non-PH), non-PH candidates also exist. However, there has not been a comprehensive comparison of these candidates. Additionally, follow-up is an important consideration for data with L-TS, but there has been limited study of the impact of follow-up time on performance of two-sample tests with L-TS. We conducted a neutral simulation study of the impact of sample size and follow-up time on type I error and power across varying effect sizes for conventional methods, methods adapted for non-PH, and a correctly-specified parametric model. When one or both groups lack L-TS, log-rank tests and one non-PH method typically have the highest power, but order varies. Surprisingly, when both groups have L-TS, these tests have non-monotonic power as a function of follow-up time, while parametric models have monotonic increasing power and the highest power at the longest follow-up time. While absolute power differs, patterns over follow-up are consistent across sample sizes. To address this for practitioners, we devise a numerical approach to predict the potential for non-monotonicity during study planning. We conclude that naïve use of conventional methods can have counterintuitive properties in settings with L-TS, and this work provides knowledge and a tool to anticipate and address these issues.