A Study of the Reliability of Agentic AI-Generated Programs
Quick summary
arXiv:2609.18298v1 Announce Type: cross Abstract: Agentic-AI based software development offers the promise of faster completion of the software, greater programmer efficiency, and more reliable code. The question is how can we verify these claims in an objective way? In this project, we attempted to answer this question based on three practices. First, we applied a typical best-practices agentic AI workflow for software development. Second, our target programs were ten well-known, release-quality human-written Linux utility programs so that we could compare the AI-generated code against a conc
Key takeaways
- arXiv:2609.18298v1 Announce Type: cross Abstract: Agentic-AI based software development offers the promise of faster completion of the software, greater programmer efficiency, and more reliable code.
- The question is how can we verify these claims in an objective way?
- In this project, we attempted to answer this question based on three practices.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments