← All papers

Extrapolation Is an Assumption: Sharp Limits for Pass@k Forecasting

Abstract

Repeated sampling can reveal capabilities and risks that are invisible in one language-model attempt, motivating forecasts of pass@$k$ from much smaller response banks. We ask what such data support without a parametric law for task difficulty. With $b$ attempts per task, the count distribution identifies exactly the first $b$ moments of the latent success-probability distribution. We prove that the worst-case diameter of pass@$k$'s identified set is exactly twice the best degree-$b$ uniform approximation error for $x^k$. A theorem of Newman and Rivlin then yields explicit binomial-tail bounds and a sharp extrapolation transition at $k\asymp b^2$. We extend the limit to finite task sets through a Hellinger modulus and to adaptive policies with a per-task cap. We also give two finite-sample confidence intervals: a global polynomial interval that attains the limiting worst-case diameter and a data-adaptive projection interval. At 128 tasks and 78 attempts per task, any honest 95% interval for population pass has expected-length lower bounds $0.276$ and $0.686$ for pass@1000 and pass@10000. In synthetic audits, narrow taskwise Beta intervals have zero coverage under seven nonparametric alternatives. Across 22 public response banks, model-free intervals include every 10,000-response reference but are necessarily wide at distant horizons. Pass@$k$ extrapolation can be useful, but its uncertainty must disclose the structural assumptions that make it possible.

Keywords: pass@k, repeated sampling, language model evaluation, extrapolation, binomial tails, forecasting

Full text (PDF) · source repository · doi:10.5281/zenodo.21698610