DeltaSelect: Affordable A/B Testing for Coding Agents
Quick summary
arXiv:2609.19607v1 Announce Type: cross Abstract: Coding-agent benchmarks are built for broad and comprehensive comparisons, not frequent development decisions. Individual runs vary, full suites are expensive, and the benchmark harness may differ from the harness used in practice. In a resampling analysis of DeepSWE's published trials, only 19.5% of tasks (22 of 113) had a fifth-percentile Pearson correlation of at least 0.50 with full-benchmark performance. The paper presents DeltaSelect, an open-source method that identifies tasks whose one-run results consistently track full-benchmark perfo
Key takeaways
- arXiv:2609.19607v1 Announce Type: cross Abstract: Coding-agent benchmarks are built for broad and comprehensive comparisons, not frequent development decisions.
- Individual runs vary, full suites are expensive, and the benchmark harness may differ from the harness used in practice.
- In a resampling analysis of DeepSWE's published trials, only 19.5% of tasks (22 of 113) had a fifth-percentile Pearson correlation of at least 0.50 with full-benchmark performance.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments