arXiv:cs.LG(机器学习,全量分类)· Alfredo Madrid-Garc\'ia, Beatriz Merino-Barbancho·· 5 小时前AI 评分36
Jev 医学基准评测:非生成式“System One”模型在医学问答上的表现
Jev in Medicine: A Benchmark Evaluation
AI 导读
研究评测了非生成式“System One”模型 Jev 1.13 在 MetaMedQA、PubMedQA、DiagnosisArena-MCQ 和 NEJM Case Challenges 四个医学基准上的表现,以 GPT-6 Sol 为参照。
正文
Abstract:Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. All 8,469 requests returned a valid answer. Jev's accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%;), lower on MetaMedQA (74.8% vs 82.7%) and much lower on DiagnosisArena-MCQ (59.8% vs 82.4%;) and the NEJM cases (61.8% vs 82.4%). On MetaMedQA, Jev's probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability of at least 0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev's probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was "I don't know or cannot answer", Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27-0.31 s; all 2,823 items cost USD 0.08. Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Task-specific validation is required before clinical use.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.34024 [cs.AI] |
| (or arXiv:2609.34024v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34024 arXiv-issued DOI via DataCite |
Submission history
From: Alfredo Madrid Garcia [view email]
[v1]
Sun, 27 Sep 2026 23:33:04 UTC (753 KB)
[v2]
Thu, 1 Oct 2026 06:03:21 UTC (753 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org