Revisiting dependence in multiple testing: empirical distribution approaches for FDP control

📅 2026-09-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对高通量数据分析中的多重假设检验问题,提出了一种基于经验累积分布函数的FDP控制方法(eFDP),无需显式建模依赖关系,通过非参数估计和多变量混合模型框架实现更准确的FDP控制。
📝 Abstract
Large-scale multiple hypothesis testing is central to the analysis of high-throughput data, where controlling false discoveries is critical. Classical procedures typically rely on theoretical null distributions and often adjust for dependence among test statistics, but these approaches may be misleading when the empirical distribution of null statistics deviates from theoretical assumptions. Motivated by this observation, we investigate the role of the empirical cumulative distribution function (c.d.f.) of null test statistics in controlling the false discovery proportion (FDP). We first show that, under an oracle scenario where the empirical c.d.f. of the test statistics for all null hypotheses is known, FDP control can be achieved optimally regardless of the dependence structure, highlighting that explicit modeling of dependence may be unnecessary. Building on this insight, we propose an empirical c.d.f.-based FDP control (eFDP) method, implemented via a multivariate mixture model framework and a nonparametric estimation procedure for the empirical c.d.f.s, establish its asymptotic convergence, and construct an FDP control procedure that achieves asymptotic FDP control. Extensive simulations demonstrate that eFDP attains more accurate FDP control and higher power than existing approaches, particularly under strong dependence, and analysis of a high-dimensional breast cancer gene expression dataset confirms its practical utility.
Problem

Research questions and friction points this paper is trying to address.

multiple hypothesis testing
false discovery proportion
empirical distribution
dependence structure
high-throughput data
Innovation

Methods, ideas, or system contributions that make the work stand out.

empirical cumulative distribution function
false discovery proportion control
multivariate mixture model