🤖 AI Summary
Existing p-hacking detection methods lack systematic evaluation of statistical power under diverse p-hacking strategies and realistic effect-size distributions. Method: We establish a theoretical mapping between p-hacking mechanisms and the resulting p-value distribution, analytically deriving power bounds for various detection approaches—particularly joint null hypothesis tests—and quantifying how publication bias modulates detection power. Contribution/Results: We demonstrate that publication bias can substantially *enhance* the power of joint tests, challenging the conventional view that it uniformly undermines detection validity. This work provides the first analytical characterization linking p-hacking structure, underlying effect-size distribution, and detection power. Empirical analysis reveals critically low power in most current applications, warning against overinterpretation of p-value distribution evidence (e.g., p-curve) as definitive proof of p-hacking.
📝 Abstract
$p$-Hacking undermines the validity of empirical studies. A flourishing empirical literature investigates the prevalence of $p$-hacking based on the distribution of $p$-values across studies. Interpreting results in this literature requires a careful understanding of the power of methods for detecting $p$-hacking. We theoretically study the implications of likely forms of $p$-hacking on the distribution of $p$-values to understand the power of tests for detecting it. Power depends crucially on the $p$-hacking strategy and the distribution of true effects. Publication bias can enhance the power for testing the joint null of no $p$-hacking and no publication bias.