arXiv Artificial Intelligence

E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

Quick summary

arXiv:2609.38335v1 Announce Type: cross Abstract: Coding agents powered by large language models (LLMs) are evolving from making localized code changes to developing complete software repositories. However, evaluating repository-scale generation remains challenging: tasks must demand system-level reasoning while ensuring that all evaluated behaviors are precisely specified and independent of any particular implementation. We introduce E2E-SWE, a benchmark for evaluating whether coding agents can build complete, functional software repositories end to end. E2E-SWE contains 186 whole-repository

Key takeaways

  • arXiv:2609.38335v1 Announce Type: cross Abstract: Coding agents powered by large language models (LLMs) are evolving from making localized code changes to developing complete software repositories.
  • However, evaluating repository-scale generation remains challenging: tasks must demand system-level reasoning while ensuring that all evaluated behaviors are precisely specified and independent of any particular implementation.
  • We introduce E2E-SWE, a benchmark for evaluating whether coding agents can build complete, functional software repositories end to end.

Why it matters

“E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗