arXiv Artificial Intelligence

VeRA: Renewing Reasoning Benchmarks with Executable Specifications

VeRA: Renewing Reasoning Benchmarks with Executable Specifications

Quick summary

arXiv:2602.13217v2 Announce Type: replace Abstract: Reasoning benchmarks need renewal along two axes: freshness and headroom. VeRA makes both executable and auditable by turning each item into a task family: a natural-language template, an input generator, and a deterministic answer program. VeRA-E draws fresh instances within a family; VeRA-H modifies the family toward harder tasks; and VeRA-H Pro selects one judge-ranked candidate from up to five validated proposals per seed. Execution checks, seed anchoring, answer discrimination, and independent human solving validate specifications and it

Key takeaways

  • arXiv:2602.13217v2 Announce Type: replace Abstract: Reasoning benchmarks need renewal along two axes: freshness and headroom.
  • VeRA makes both executable and auditable by turning each item into a task family: a natural-language template, an input generator, and a deterministic answer program.
  • VeRA-E draws fresh instances within a family; VeRA-H modifies the family toward harder tasks; and VeRA-H Pro selects one judge-ranked candidate from up to five validated proposals per seed.

Why it matters

This model development creates a new option for users and a new testing obligation for developers. A fixed evaluation set comparing quality, cost and failure behavior is more useful than launch claims.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗