Code Owns the Simulation, Jev Owns the Evaluation
Quick summary
arXiv:2610.01834v1 Announce Type: new Abstract: Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the input describes, which we call \emph{evaluation}. Specifically, it solves 99\% of the counterintuitive Cogni
Key takeaways
- arXiv:2610.01834v1 Announce Type: new Abstract: Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option.
- This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with.
- We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary.
Why it matters
“Code Owns the Simulation, Jev Owns the Evaluation” highlights the need for repeatable measurement rather than a single impressive demonstration. Independent validation across datasets and clearly stated limitations determine whether a result can guide product decisions.

Member comments