arXiv Artificial Intelligence

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Quick summary

arXiv:2609.26550v3 Announce Type: replace Abstract: LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict mus

Key takeaways

  • arXiv:2609.26550v3 Announce Type: replace Abstract: LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly.
  • We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge.
  • Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict mus

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗