What Fixed-Rollout pass@k Evaluations Can Identify

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究探讨了固定采样次数下pass@k评估方法的局限性,通过统计模型分析指出该方法在超出采样数时无法准确识别性能指标,并提出了新的评估标准。
📝 Abstract
Repeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples collected per problem. We show that, in the pooled/random-task conditional-Binomial model, fixed-n success counts identify only the n free moments of the latent per-task success distribution. Consequently, direct pass@k is identified for k <= n, but generic extrapolated pass@k, tail exponents, and tail constants are not identified for k > n, even with arbitrarily many exchangeable tasks at the same rollout budget. This is stronger than the observation that the usual estimator is undefined beyond n: it characterizes the information missing from the fixed-depth count-law experiment. We give exact count-law-preserving constructions with incompatible extrapolations, state the exceptional unique-extension case, and compute sharp population identified intervals through Hausdorff principal representations. On the public 10,000-rollout-per-problem release of Brown et al., counterfactual n = 16 evaluations leave failure at k = 1000 ambiguous by factors from 1.5 to over 2,600 across four MATH/GSM8K/CodeContests configurations. The calibration shows that intermediate-scale failure share alone does not determine width. Our result does not reject parametric inference-time scaling laws; it supplies the nonparametric baseline against which their assumptions can be evaluated. We give an exact, conservative one-coordinate finite-task confidence certificate and a reporting standard separating direct estimates, identified sets, and model-conditioned forecasts.
Problem

Research questions and friction points this paper is trying to address.

fixed-rollout
pass@k
extrapolation
sample limitation
latent distribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

fixed-n success counts
latent per-task success distribution
extrapolated pass@k
count-law-preserving constructions
finite-task confidence certificate
🔎 Similar Papers
No similar papers found.
Pranav Singh
Pranav Singh
New York University
Representation LearningDeep LearningComputer VisionMedical Image Analysis
P
Prashant Singh
Department of Mathematics, Indian Institute of Technology Ropar, Ropar, Punjab 140001, India