arXiv Artificial Intelligence

From Verification Failures to Reusable Guidance for Coding Agents

From Verification Failures to Reusable Guidance for Coding Agents

Quick summary

arXiv:2609.39022v1 Announce Type: cross Abstract: Coding agents need to establish that a program satisfies a specification and that the specification captures the requested behavior. We study how expert diagnosis of verification failures can become reusable guidance for this work. Our approach combines executable language definitions in the K framework with a kit of procedures for constructing specifications, repairing proofs, and auditing their adequacy. A human-guided development campaign on HumanEval, a benchmark of 164 Python programming tasks, achieves a 164/164 success rate with the sema

Key takeaways

  • arXiv:2609.39022v1 Announce Type: cross Abstract: Coding agents need to establish that a program satisfies a specification and that the specification captures the requested behavior.
  • We study how expert diagnosis of verification failures can become reusable guidance for this work.
  • Our approach combines executable language definitions in the K framework with a kit of procedures for constructing specifications, repairing proofs, and auditing their adequacy.

Why it matters

“From Verification Failures to Reusable Guidance for Coding Agents” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗