Score
Formulating anytime-valid sequential hypothesis tests (e-processes) with uniform type-I error control (e.g., via Ville's inequality and calibration) and methods to incorporate priors or detect cheating while preserving error guarantees.
Existing e-BH procedures lack order-invariance over e-processes, causing test conclusions to reverse spuriously upon addition of irrelevant data and failing to control the false discovery rate (FDR) — or even the family-wise error rate (FWER) — under arbitrary dependence. This paper provides the first rigorous proof that e-BH violates FDR control in this setting. Method: We propose a novel, order-invariant multiple testing framework built on e-process upper bounds, featuring a dependence-structure-adaptive calibrator. Contribution/Results: Our method guarantees strict FDR control at level α (i.e., FDR-sup ≤ α) for arbitrary dependence structures among hypotheses. It eliminates temporal instability in rejection sets induced by sequential data arrival, ensuring robustness and reproducibility in dynamic data environments. Theoretical guarantees are established without restrictive assumptions on dependence, and the procedure is computationally tractable.
This work addresses the problem of real-time, data-adaptive lower bounding of the number of true discoveries in online multiple testing, under the constraint that decisions must be based solely on past hypotheses and data—without lookahead. To tackle this challenge, we first establish that admissibility of online closed testing is equivalent to the availability of valid e-values at each time step. Building on this equivalence, we propose a novel online closed testing framework grounded in products of e-values, accompanied by an efficient algorithm and new e-value hedging and boosting mechanisms to enhance statistical power. Our method unifies modeling of exchangeable and arbitrarily dependent test statistics, ensuring strict control of both false discovery rate (FDR) and false discovery exceedance (FDX) under arbitrary dependence, while delivering tight lower confidence bounds on the number of true discoveries. Compared to state-of-the-art approaches, our method achieves theoretical unification and improved performance, enabling the first online procedure for true discovery control that simultaneously delivers strong empirical power and rigorous finite-sample guarantees.
Classical Benjamini–Hochberg (BH) and e-BH procedures—designed for offline multiple testing—are not directly applicable to streaming data, and existing online FDR control methods lack rigorous guarantees for stopping-time FDR. Method: We propose an online BH/e-BH framework incorporating a *delayed rejection* mechanism. Contribution/Results: Our method is the first to simultaneously achieve theoretical consistency—provably controlling both fixed-time and stopping-time FDR under arbitrary dependence—and reduction—recovering standard offline BH/e-BH as special cases. We unify and extend the analysis of major online procedures, proving their stopping-time FDR control via sequential decision theory, dependence-agnostic bounding, and stopping-time extension techniques. Crucially, our framework attains the *same* FDR bound as offline BH/e-BH. An open-source implementation enables real-time streaming hypothesis testing.
This paper addresses the problem of testing exchangeability of random variables and invariance under compact groups. Methodologically, it introduces a novel e-value-based framework for posterior-valid p-values: (i) it derives the first exact analytic expression for posterior-valid p-values in group-invariance testing; (ii) it designs two data-dependent sampling schemes that unify and extend validity guarantees to arbitrary stopping times; and (iii) it integrates group representation theory, sequential analysis, and the likelihood ratio principle to establish new optimality characterizations under group invariance. Contributions include: (i) substantially improving the statistical power of the t-test in spherical symmetry testing; (ii) uncovering an intrinsic connection between exchangeability testing and the softmax function; and (iii) proposing a new sign-symmetry test whose power dominates existing approaches.
Wald’s sequential probability ratio test (SPRT) suffers from overshoot at stopping times, preventing approximate thresholds—such as ((1-eta)/alpha) and (eta/(1-alpha))—from strictly controlling Type I/II error rates (when (eta > 0)) or guaranteeing optimality (when (eta = 0)). This paper introduces “sequential boosting”, a novel method that eliminates overshoot by constructing a corrected likelihood ratio statistic. It achieves, for the first time: (1) exact (alpha)/(eta) error control with strictly smaller expected sample size than approximate SPRT when (eta > 0); (2) optimal power-one performance—matching the theoretical lower bound on expected sample size—when (eta = 0); and (3) natural generalizations to confidence sequences, sampling-without-replacement settings, and conformal martingale frameworks. Theoretical analysis proves precise error calibration, while simulations demonstrate substantial sample-size reduction. The method is plug-and-play and broadly applicable across sequential inference paradigms.
This work addresses the problem of multiple hypothesis testing for edge distributions across multiple data streams. It proposes a sequential testing procedure that, for the first time, systematically incorporates arbitrary forms of prior information about the configuration of true and false hypotheses—such as known values or lower bounds on the number of active streams under each hypothesis, or mutual exclusivity constraints—while rigorously controlling the familywise error rate. By integrating sequential analysis with a search strategy over minimal alternative hypothesis configurations, the method achieves asymptotic optimality in terms of expected sample size among all valid procedures, without compromising reliability. Theoretical analysis establishes its computational efficiency and asymptotic optimality, and numerical experiments further demonstrate its substantial advantages in both testing efficiency and accuracy.
This work addresses the problem of constructing asymptotically log-optimal e-processes from asymptotically optimal sequential tests. To this end, it introduces a novel class of WAIT e-processes (Weighted Average of Stopped Indicator Tests), which for the first time enables the reverse construction of asymptotically log-optimal e-processes from asymptotically optimal sequential tests. The proposed framework not only establishes a bidirectional equivalence between these two notions of optimality—thereby completing a key theoretical gap—but also clarifies subtle distinctions among different definitions of asymptotic optimality in the literature. The resulting WAIT e-processes are shown to grow to infinity at the optimal rate under the alternative hypothesis, thereby verifying their log-optimality.
This work addresses the challenge in sequential hypothesis testing where model misspecification or estimation error prevents exact construction of e-variables, thereby lacking finite-sample guarantees. We introduce, for the first time, the notion of an asymptotic e-process, defined as a doubly indexed stochastic process $(E_{m,n})$, whose limiting behavior as $m \to \infty$ approximates a standard e-process. We establish its connection to asymptotic supermartingales, derive a corresponding variant of Ville’s inequality, and provide practical construction methods. This framework unifies the theoretical foundation for approximate e-variables, offering sequential inference guarantees under controllable approximation error and explicitly quantifying the trade-off between approximation accuracy and the effective monitoring horizon $r_m$.
This study addresses a central challenge in statistical inference for clinical trials: achieving high operational flexibility—such as sample size re-estimation and treatment selection—while rigorously controlling the Type I error rate. The authors systematically integrate confirmatory adaptive designs with e-value-based anytime-valid testing methods, establishing for the first time their formal equivalence through conditional error functions and combination tests, while clarifying their distinct emphases on flexibility. The work constructs a theoretical bridge between these two frameworks: the e-value paradigm enhances optional continuation and loss control, whereas adaptive design principles can refine e-value testing strategies. This synthesis lays both a theoretical foundation and a practical pathway for developing next-generation inferential methods that simultaneously ensure strict error control and substantial procedural adaptability.
This study addresses the problem of determining whether high-frequency monitoring data return to their pre-intervention baseline distribution following an intervention. The authors propose a sequential testing procedure that requires no assumptions about the underlying data distribution. The method constructs a discrepancy measure via universal inference and combines it with individualized empirical calibration to form a non-negative supermartingale, yielding an e-process that enables valid detection of the recovery time at any arbitrary stopping point without specifying a null model. Theoretical analysis provides finite-sample bounds on the calibration error, and both simulations and a clinical case study demonstrate the method’s superior performance in accurately identifying the time at which baseline conditions are restored.