π€ AI Summary
This study addresses critical limitations in current online A/B testing methodologies, where fixed-horizon tests suffer from inflated Type I error due to repeated monitoring, and prevailing sequential approaches struggle to simultaneously support futility stopping, control Type II error, and respect minimum detectable effect constraints. To overcome these challenges, the authors propose SPRT-z, a novel procedure grounded in Waldβs sequential probability ratio test. SPRT-z leverages large-sample normal approximations for computational efficiency, incorporates a scale-free horizon calibration (SFHC) mechanism to preserve statistical power under discrete monitoring, and employs a median-unbiased estimator derived from Brownian motion to correct bias induced by early stopping. Empirical evaluations demonstrate that the method rigorously controls both Type I and Type II errors, substantially reduces required sample sizes, and yields confidence intervals with coverage probabilities closely matching nominal levels.
π Abstract
Modern online experimentation platforms produce data at scale and continuously. However, practitioners routinely apply Fixed Horizon Testing (FHT) under repeated peeking, inflating Type I error and reducing decision quality. Popular always valid sequential methods control Type I error under peeking and enable early stopping for efficacy, but do not natively support early futility stopping, launch criteria tied to a business-relevant minimum detectable effect, or Type II error control. As an alternative that satisfies these useful properties, we revive Wald's Sequential Probability Ratio Test (SPRT) for online experimentation with three novel contributions: (1) SPRT-z, an adaptation of Hajnal's sequential $t$-test, leverages large sample normal approximation to eliminate computational bottlenecks inherent to the scale of modern A/B tests and enables the Brownian motion-based methods used in (2) and (3); (2) Scale-Free Horizon Calibration (SFHC) is a Monte Carlo bisection procedure on the standardised $Z$-scale that sets a maximum sample size preserving nominal power under discrete monitoring with futility stopping; (3) A Brownian Median Unbiased Estimator and accompanying confidence intervals correct the upward bias induced by early stopping across all stopping regions via a six-region stagewise ordering of the sample space. A simulation study shows this workflow appropriately controls Type I and II error, reduces sample size relative to FHT, and ameliorates estimation bias from early stopping with close-to-nominal confidence interval coverage in most scenarios studied.