arXiv Artificial Intelligence

Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation

Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation

Quick summary

arXiv:2608.29463v1 Announce Type: cross Abstract: A benchmark score is a joint property of the model, the evaluation harness, the elicitation budget, the sampled population, and contamination status. Leaderboards publish the model and the score, so capability and leakage stay observationally equivalent. Existing taxonomies classify contamination for automated detection, not the question a reporter faces at publication: given the mitigations already applied, which validity threats remain open? We introduce a taxonomy organized by the mitigation each type defeats -- direct, derivative, temporal,

Key takeaways

  • arXiv:2608.29463v1 Announce Type: cross Abstract: A benchmark score is a joint property of the model, the evaluation harness, the elicitation budget, the sampled population, and contamination status.
  • Leaderboards publish the model and the score, so capability and leakage stay observationally equivalent.
  • Existing taxonomies classify contamination for automated detection, not the question a reporter faces at publication: given the mitigations already applied, which validity threats remain open?

Why it matters

“Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗