When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

📅 2026-08-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fundamental ambiguity in existing contamination audits, which often conflate truly clean data with insufficient detection power. To resolve this, the authors propose a theoretical framework based on sparse mixture models of the form \( Q_\alpha = (1-\alpha)P_0 + \alpha P_1 \), leveraging control-group estimation to assess statistical power. By integrating sample-splitting certificates with a two-stage planner, the method enables calibrated and reliable auditing. Key contributions include establishing distribution-free lower-bound certificates for contamination proportion, uncovering the failure mechanism of Gaussian budget calibration in small-sample regimes, and providing a corrective solution. Empirical results demonstrate high predictive accuracy of power curves (\( R^2 = 0.83\text{–}0.98 \)) across six channels and successfully reproduce the sensitivity ranking of injected contaminations: verbatim > paraphrase > surface. The repaired budgeting scheme proves conservatively effective, and non-rejection conclusions require joint evaluation of power, budget, and validity gating.
📝 Abstract
Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power. We formalize this distinction for a benchmark in which an unknown fraction alpha of items was seen during training. With matched clean and seen controls, the behavioral channel is the sparse mixture Q_alpha = (1 - alpha) P_0 + alpha P_1, and an exact second-moment argument shows that detectability is governed by alpha * rho * sqrt(m), where rho^2 = chi^2(P_1 || P_0) measures behavioral separability. Any scalar detector reduces to its efficacy, ef = |E_1 f - E_0 f| / sqrt(Var_0(f)) <= rho, which can be estimated from controls before the audit is run. A separate sample-split certificate lower-bounds alpha distribution-free, without requiring an orientation assumption. Our empirical finding is two-sided. Frozen calibration efficacy predicts held-out power curves, with R^2 = 0.83-0.98 across six exact-permutation channels, but the efficacy-only Gaussian budget is miscalibrated at the small sample sizes it prescribes, failing in 9/9 gate-passing channels even though efficacy itself transports. The failure is in the inversion, not the calibration. A predeclared two-stage planner that simulates the deployed test repairs the budgets, is uniformly conservative, and abstains when its probe does not transport. The certificate is valid but vacuous at audit scale, and a five-seed paired injection study recovers the mechanism ordering verbatim > paraphrase > surface, in which the apparent answer-only signal is explained by baseline drift. We report the audit contract and its failures together: a non-rejection is interpretable only alongside the efficacy, budget, and validity gates that produced it.
Problem

Research questions and friction points this paper is trying to address.

benchmark contamination
detectability
behavioral auditing
training data leakage
statistical power
Innovation

Methods, ideas, or system contributions that make the work stand out.

benchmark contamination
statistical power calibration
efficacy estimation
distribution-free certificate
two-stage audit planner