π€ AI Summary
This work addresses the problem of real-time, data-adaptive lower bounding of the number of true discoveries in online multiple testing, under the constraint that decisions must be based solely on past hypotheses and dataβwithout lookahead. To tackle this challenge, we first establish that admissibility of online closed testing is equivalent to the availability of valid e-values at each time step. Building on this equivalence, we propose a novel online closed testing framework grounded in products of e-values, accompanied by an efficient algorithm and new e-value hedging and boosting mechanisms to enhance statistical power. Our method unifies modeling of exchangeable and arbitrarily dependent test statistics, ensuring strict control of both false discovery rate (FDR) and false discovery exceedance (FDX) under arbitrary dependence, while delivering tight lower confidence bounds on the number of true discoveries. Compared to state-of-the-art approaches, our method achieves theoretical unification and improved performance, enabling the first online procedure for true discovery control that simultaneously delivers strong empirical power and rigorous finite-sample guarantees.
π Abstract
In contemporary research, data scientists often test an infinite sequence of hypotheses $H_1,H_2,ldots $ one by one, and are required to make real-time decisions without knowing the future hypotheses or data. In this paper, we consider such an online multiple testing problem with the goal of providing simultaneous lower bounds for the number of true discoveries in data-adaptively chosen rejection sets. In offline multiple testing, it has been recently established that such simultaneous inference is admissible iff it proceeds through (offline) closed testing. We establish an analogous result in this paper using the recent online closure principle. In particular, we show that it is necessary to use an anytime-valid test for each intersection hypothesis. This connects two distinct branches of the literature: online testing of multiple hypotheses (where the hypotheses appear online), and sequential anytime-valid testing of a single hypothesis (where the data for a fixed hypothesis appears online). Motivated by this result, we construct a new online closed testing procedure and a corresponding short-cut with a true discovery guarantee based on multiplying sequential e-values. This general but simple procedure gives uniform improvements over the state-of-the-art methods but also allows to construct entirely new and powerful procedures. In addition, we introduce new ideas for hedging and boosting of sequential e-values that provably increase power. Finally, we also propose the first online true discovery procedures for exchangeable and arbitrarily dependent e-values.