arXiv Artificial Intelligence

Evaluating Exact Output and Checkpoint-State Prediction in Real Programs

Evaluating Exact Output and Checkpoint-State Prediction in Real Programs

Quick summary

arXiv:2610.11889v1 Announce Type: cross Abstract: We present a benchmark for predicting final output and checkpoint state from source and input alone. It extends CRUXEval-style output prediction with paired shorter- and longer-trace inputs and checkpoints inside and after a loop. The benchmark contains 400 cases from 371 Python and C++ programs, evaluated under seven settings from four model families without tools or code execution. Of 11,200 planned predictions, 11,151 produced gradable responses. Reasoning-enabled settings outperform their off counterparts by 33.1 to 55.2 percentage points o

Key takeaways

  • arXiv:2610.11889v1 Announce Type: cross Abstract: We present a benchmark for predicting final output and checkpoint state from source and input alone.
  • It extends CRUXEval-style output prediction with paired shorter- and longer-trace inputs and checkpoints inside and after a loop.
  • The benchmark contains 400 cases from 371 Python and C++ programs, evaluated under seven settings from four model families without tools or code execution.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗