arXiv Artificial Intelligence

What Fixed-Rollout pass@k Evaluations Can Identify

What Fixed-Rollout pass@k Evaluations Can Identify

Quick summary

arXiv:2609.09245v1 Announce Type: cross Abstract: Repeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples collected per problem. We show that, in the pooled/random-task conditional-Binomial model, fixed-n success counts identify only the n free moments of the latent per-task success distribution. Consequently, direct pass@k is identified for k n, even with arbitrarily many exchangeable tasks at the same rollout budget. This is stronger than the observation that the usual estimator is undefined beyond n: it characterizes the information missing from the

Key takeaways

  • arXiv:2609.09245v1 Announce Type: cross Abstract: Repeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples collected per problem.
  • We show that, in the pooled/random-task conditional-Binomial model, fixed-n success counts identify only the n free moments of the latent per-task success distribution.
  • Consequently, direct pass@k is identified for k n, even with arbitrarily many exchangeable tasks at the same rollout budget.

Why it matters

“What Fixed-Rollout pass@k Evaluations Can Identify” should be evaluated beyond branding and benchmark scores. Its practical importance will emerge in task accuracy, latency, unit cost, safety and integration with real workflows.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗