local power analysis

Theoretical and empirical analysis of statistical-test detectability under local alternatives, contiguous deviations, and geometry-induced drift (e.g., RKHS effects), including analytic characterizations of local consistency and methods for combining p-values while assessing power.

localpoweranalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.31
Aug 01, 2026Aug 01, 2026
Career
Value
No comparison yet
$191K/year
Aug 01, 2026Aug 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

On the Robustness of Kernel Goodness-of-Fit Tests

Aug 11, 2024
XL
Xing Liu
🏛️ Imperial College London | University College London

Existing kernel-based goodness-of-fit (GoF) tests fail under both qualitative and quantitative robustness, and robustification strategies—such as tilted kernels—cannot simultaneously satisfy both criteria in GoF testing. Method: Addressing the practical question “Is the model sufficiently accurate?”, we propose the first robust GoF testing framework based on a kernel Stein discrepancy (KSD) ball. This framework rigorously formalizes robust GoF testing and theoretically establishes that conventional kernel tests—and their tilted-kernel variants—lack dual robustness. Contribution/Results: Our test achieves both stability and statistical power under diverse contamination models—including Huber contamination and density bands—enabling unified robust modeling. Empirical evaluation confirms its effectiveness in finite-sample settings, resolving a long-standing challenge in designing robust kernel-based GoF tests.

Addressing model adequacy under data perturbationsDeveloping robust tests using kernel Stein discrepancyTesting robustness of kernel goodness-of-fit methods

This study addresses the long-standing problem of characterizing the existence of nontrivial (strictly unbiased) hypothesis tests in the absence of a common dominating measure. By examining the closure of the convex hull of a set of probability measures within the space of bounded finitely additive measures, and leveraging topological separation properties under the total variation distance, the authors establish necessary and sufficient conditions for testability without requiring a reference control measure. This work completes the theoretical program initiated by Le Cam, providing the first complete characterization of testability in full generality. Illustrative examples further highlight the essential roles played by measure-theoretic and convex-analytic considerations in this foundational result.

convex hullsfinitely additive measureshypothesis testing

Practical Kernel Tests of Conditional Independence

Feb 20, 2024
RP
Roman Pogodin
🏛️ McGill University | Mila | Gatsby Computational Neuroscience Unit | University College London | University of British Columbia | Amii

Addressing the dual challenges of inflated Type I error rates (loss of test-level control) and low statistical power in conditional independence testing, this paper proposes a data-efficient kernel-based testing framework. The method employs kernel ridge regression and introduces, for the first time in this setting, three principled bias-correction strategies: data splitting, auxiliary data utilization, and restriction to simplified function classes—ensuring rigorous asymptotic and finite-sample control of the significance level. Theoretically, the approach guarantees convergence of the Type I error rate to the nominal significance level while enhancing detection power for complex dependency structures. Extensive experiments on diverse synthetic and real-world datasets demonstrate that the proposed method achieves precise Type I error control and substantially outperforms state-of-the-art competitors—including KCIT and RCIT—in statistical power, with improved robustness and reliability.

Addressing bias in test statistics from kernel ridge regression methodsDeveloping kernel-based tests for conditional independence with accurate false positive controlImproving test level accuracy while maintaining competitive statistical power

The Power of Tests for Detecting $p$-Hacking

May 16, 2022
GE
G. Elliott
🏛️ University of California San Diego | Queen’s University | University of Michigan

Existing p-hacking detection methods lack systematic evaluation of statistical power under diverse p-hacking strategies and realistic effect-size distributions. Method: We establish a theoretical mapping between p-hacking mechanisms and the resulting p-value distribution, analytically deriving power bounds for various detection approaches—particularly joint null hypothesis tests—and quantifying how publication bias modulates detection power. Contribution/Results: We demonstrate that publication bias can substantially *enhance* the power of joint tests, challenging the conventional view that it uniformly undermines detection validity. This work provides the first analytical characterization linking p-hacking structure, underlying effect-size distribution, and detection power. Empirical analysis reveals critically low power in most current applications, warning against overinterpretation of p-value distribution evidence (e.g., p-curve) as definitive proof of p-hacking.

Analyze p-value distributions under various p-hacking strategiesIdentify most effective tests for detecting p-hackingStudy power of tests detecting p-hacking in research

Post-hoc and Anytime Valid Permutation and Group Invariance Testing

Oct 02, 2023
NW
Nick W. Koning
🏛️ Erasmus University Rotterdam

This paper addresses the problem of testing exchangeability of random variables and invariance under compact groups. Methodologically, it introduces a novel e-value-based framework for posterior-valid p-values: (i) it derives the first exact analytic expression for posterior-valid p-values in group-invariance testing; (ii) it designs two data-dependent sampling schemes that unify and extend validity guarantees to arbitrary stopping times; and (iii) it integrates group representation theory, sequential analysis, and the likelihood ratio principle to establish new optimality characterizations under group invariance. Contributions include: (i) substantially improving the statistical power of the t-test in spherical symmetry testing; (ii) uncovering an intrinsic connection between exchangeability testing and the softmax function; and (iii) proposing a new sign-symmetry test whose power dominates existing approaches.

Design optimal e-values for group invariance with utility targetsGeneralize rank- and sign-based testing to compact groupsQuantify evidence against exchangeability and group invariance using e-values

Latest Papers

What's happening recently
View more

This study addresses the deviation of empirical probability integral transforms (PIT) from the theoretical uniform distribution under finite samples, a phenomenon induced by the two-stage sampling structure that invalidates conventional one-sample uniformity tests. The work systematically demonstrates that this non-uniformity arises from dependence structures and variance distortions introduced either by a reference sample or a rolling window: the former converges to a two-sample Kolmogorov–Smirnov distribution, while the latter exhibits temporal autocorrelation. Building on probability integral transform theory, empirical quantile estimation, and two-sample KS asymptotics, this paper establishes—for the first time—that empirical PIT values cannot be treated as independent uniform random variables. Leveraging these insights, the authors develop a corrected statistical inference framework specifically tailored for backtesting forecast calibration.

backtestingdependenceempirical p-values

This work addresses the challenge of conformal anomaly detection under data distribution shifts and limited sample regimes, where standard importance weighting drastically reduces the effective sample size, leading to overly conservative p-values or inflated variance that impairs anomaly identification. To overcome this, the authors propose a continuous inference relaxation framework that, for the first time, integrates continuous weighted kernel density estimation into conformal anomaly detection. By locally adapting to non-stationary data distributions, the method decouples the trade-off between tail resolution and stability while preserving marginal coverage guarantees. This approach mitigates the loss of statistical power induced by discretization, eliminates Monte Carlo variability, and substantially enhances detection capability in low-data settings, successfully recovering anomalies missed by discrete baseline methods.

conformal anomaly detectiondistribution shiftimportance weighting

This study addresses the problem of testing conditional independence between random variables \( X \) and \( Y \) given a confounding variable \( Z \). It proposes a local permutation test based on data-adaptive binning—such as equal-count binning—where permutations of \( X \) and \( Y \) are performed within each subregion defined by \( Z \). The method provides, for the first time, finite-sample Type I error control guarantees for arbitrary test statistics. Under linear confounding models, it achieves power comparable to that of the oracle likelihood ratio test. Theoretical analysis shows that a constant bin size suffices to attain performance on par with increasing bin sizes, and numerical experiments confirm the method’s statistical efficiency and practical utility.

conditional independenceconfounderdata-adaptive binning

This work addresses a critical limitation in existing multiple testing procedures, which control only the expected false discovery proportion (FDP) and lack high-probability guarantees for the realized FDP, particularly when data-driven thresholds are employed, thereby compromising statistical validity. The authors propose a distribution-free, finite-sample valid framework that constructs a high-probability simultaneous envelope around the empirical distribution function of conformal p-values under the null hypothesis. This approach yields, for the first time, a uniform high-probability upper bound on the FDP that holds simultaneously over all possible rejection thresholds. The method accommodates arbitrary post-hoc threshold selection and allows users to tailor the envelope’s shape to obtain tighter bounds in regions of interest. Empirical evaluations on both synthetic and real-world data demonstrate that the resulting bounds are not only valid but also substantially less conservative than those from existing methods.

Conformal InferenceDistribution-Free BoundsFalse Discovery Proportion

This study addresses the hypothesis testing problem of detecting a hidden geometric subgraph induced by random points on a sphere within an Erdős–Rényi random graph. By integrating information-theoretic lower bounds, random geometric graph models, and the low-degree polynomial algorithmic framework, the work rigorously establishes—for the first time—a sharp “easy–hard–impossible” three-phase transition structure for this detection task. It precisely characterizes the statistical detectability threshold and constructs an algorithm that achieves this bound. Furthermore, leveraging the low-degree polynomial method, the paper provides the first proof of a strict computational-to-statistical gap, demonstrating that while detection is statistically possible beyond a certain threshold, no polynomial-time algorithm can succeed there, thereby delineating fundamental limits of computational efficiency.

computational complexityhidden geometryhypothesis testing

Hot Scholars

SL

Subhash Lakshminarayana

University of Warwick, School of Engineering
Cyber-Physical system securityWireless Communications
ZX

Zhiyao Xie

Assistant Professor, Hong Kong University of Science and Technology
EDAMachine learningVLSI circuits and systems
ML

Mengming Li

Hong Kong University of Science and Technology
Computer Architecture
HD

Hans D. Schotten

Univ. of Kaiserslautern, RPTU Kaiserslautern, DFKI GmbH
Mobile and wireless communicationsindustrial radioindustrial internetsecurity
SR

Suman Rath

The University of Tulsa
Energy SystemsCybersecurityArtificial Intelligence