arXiv Artificial Intelligence

ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance

ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance

Quick summary

arXiv:2609.09458v1 Announce Type: new Abstract: As LLM agents move from answering questions to carrying out procedures, failures can be unwarranted rather than visibly wrong: the final response looks acceptable even though the system skipped the check, branch, dependency, or invariant that made the answer justified. Output-only evaluation sees the answer, and trace-aware judging sees activity, but neither identifies which obligations were active for the query. We introduce CONTRACTEVAL, a diagnostic framework for making those active obligations explicit. It represents procedural instructions a

Key takeaways

  • arXiv:2609.09458v1 Announce Type: new Abstract: As LLM agents move from answering questions to carrying out procedures, failures can be unwarranted rather than visibly wrong: the final response looks acceptable even though the system skipped the check, branch, dependency, or invariant that made the answer justified.
  • Output-only evaluation sees the answer, and trace-aware judging sees activity, but neither identifies which obligations were active for the query.
  • We introduce CONTRACTEVAL, a diagnostic framework for making those active obligations explicit.

Why it matters

“ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗