What Fixed-Rollout pass@k Evaluations Can Identify
Quick summary
arXiv:2609.09245v1 Announce Type: cross Abstract: Repeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples collected per problem. We show that, in the pooled/random-task conditional-Binomial model, fixed-n success counts identify only the n free moments of the latent per-task success distribution. Consequently, direct pass@k is identified for k n, even with arbitrarily many exchangeable tasks at the same rollout budget. This is stronger than the observation that the usual estimator is undefined beyond n: it characterizes the information missing from the
Key takeaways
- arXiv:2609.09245v1 Announce Type: cross Abstract: Repeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples collected per problem.
- We show that, in the pooled/random-task conditional-Binomial model, fixed-n success counts identify only the n free moments of the latent per-task success distribution.
- Consequently, direct pass@k is identified for k n, even with arbitrarily many exchangeable tasks at the same rollout budget.
Why it matters
“What Fixed-Rollout pass@k Evaluations Can Identify” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Member comments