Jev in Medicine: A Benchmark Evaluation
Quick summary
arXiv:2609.34024v2 Announce Type: replace Abstract: Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognitio
Key takeaways
- arXiv:2609.34024v2 Announce Type: replace Abstract: Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them.
- Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown.
- We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges.
Why it matters
The value of this work lies as much in how it was tested as in the claim itself. Sample design, baselines, uncertainty and replication help separate a laboratory result from real-world impact.

Member comments