arXiv Artificial Intelligence

Code Owns the Simulation, Jev Owns the Evaluation

Code Owns the Simulation, Jev Owns the Evaluation

Quick summary

arXiv:2610.01834v1 Announce Type: new Abstract: Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the input describes, which we call \emph{evaluation}. Specifically, it solves 99\% of the counterintuitive Cogni

Key takeaways

  • arXiv:2610.01834v1 Announce Type: new Abstract: Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option.
  • This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with.
  • We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary.

Why it matters

“Code Owns the Simulation, Jev Owns the Evaluation” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗