arXiv Artificial Intelligence

An Empirical Study of Harness Design for Coding Agents

An Empirical Study of Harness Design for Coding Agents

Quick summary

arXiv:2609.20804v1 Announce Type: new Abstract: Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management. Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we e

Key takeaways

  • arXiv:2609.20804v1 Announce Type: new Abstract: Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear.
  • To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management.
  • Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we e

Why it matters

“An Empirical Study of Harness Design for Coding Agents” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗