WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses
Quick summary
arXiv:2609.36635v1 Announce Type: cross Abstract: Bug validation asks a coding agent to produce an executable witness for a reported bug. The witness combines a concrete input with a testing harness and exposes faulty behavior during execution. Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs and witnesses or require manual construction. We present WitnessGym, an automated framework for constructing bug-validation benchmarks through bug injection. It injects bugs into test-reached paths of real projects, rebuilds each
Key takeaways
- arXiv:2609.36635v1 Announce Type: cross Abstract: Bug validation asks a coding agent to produce an executable witness for a reported bug.
- The witness combines a concrete input with a testing harness and exposes faulty behavior during execution.
- Such evidence makes audit findings actionable, yet benchmark evaluation is difficult when cases reuse public historical bugs and witnesses or require manual construction.
Why it matters
“WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments