RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?
Quick summary
arXiv:2609.28850v1 Announce Type: new Abstract: Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments, work that AI agents increasingly do. We introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences. For each paper we fix in advance the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget. An agent must reproduce that result using the paper and whatever its authors released. What the authors released decides the difficulty tier. Run-tier r
Key takeaways
- arXiv:2609.28850v1 Announce Type: new Abstract: Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments, work that AI agents increasingly do.
- We introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences.
- For each paper we fix in advance the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget.
Why it matters
“RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?” exposes the compute, energy and supply-chain layer behind model competition. Capacity shifts can influence model costs, service availability and the ability of smaller companies to compete.

Member comments