arXiv Artificial Intelligence

Jev in Medicine: A Benchmark Evaluation

Jev in Medicine: A Benchmark Evaluation

Quick summary

arXiv:2609.34024v2 Announce Type: replace Abstract: Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognitio

Key takeaways

  • arXiv:2609.34024v2 Announce Type: replace Abstract: Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them.
  • Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown.
  • We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges.

Why it matters

The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Kaynak sitede devamını oku: arXiv Artificial Intelligence ↗