arXiv Artificial Intelligence

Can Coding Agents Reproduce Findings in Computational Materials Science?

Can Coding Agents Reproduce Findings in Computational Materials Science?

Quick summary

arXiv:2605.00803v2 Announce Type: replace-cross Abstract: Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only strong coding ability, but also the ability to navigate complex, domain-specific procedures and to interpret results in the context of scientific claims. To address this question, we present AutoMat, a benchmark for evaluating LLM-based agents' ability to repr

Key takeaways

  • arXiv:2605.00803v2 Announce Type: replace-cross Abstract: Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks.
  • However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only strong coding ability, but also the ability to navigate complex, domain-specific procedures and to interpret results in the context of scientific claims.
  • To address this question, we present AutoMat, a benchmark for evaluating LLM-based agents' ability to repr

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗