arXiv Artificial Intelligence

DeltaSelect: Affordable A/B Testing for Coding Agents

DeltaSelect: Affordable A/B Testing for Coding Agents

Quick summary

arXiv:2609.19607v1 Announce Type: cross Abstract: Coding-agent benchmarks are built for broad and comprehensive comparisons, not frequent development decisions. Individual runs vary, full suites are expensive, and the benchmark harness may differ from the harness used in practice. In a resampling analysis of DeepSWE's published trials, only 19.5% of tasks (22 of 113) had a fifth-percentile Pearson correlation of at least 0.50 with full-benchmark performance. The paper presents DeltaSelect, an open-source method that identifies tasks whose one-run results consistently track full-benchmark perfo

Key takeaways

  • arXiv:2609.19607v1 Announce Type: cross Abstract: Coding-agent benchmarks are built for broad and comprehensive comparisons, not frequent development decisions.
  • Individual runs vary, full suites are expensive, and the benchmark harness may differ from the harness used in practice.
  • In a resampling analysis of DeepSWE's published trials, only 19.5% of tasks (22 of 113) had a fifth-percentile Pearson correlation of at least 0.50 with full-benchmark performance.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗