跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Alfredo Madrid-Garc\'ia, Beatriz Merino-Barbancho·· 5 小时前AI 评分36

Jev 医学基准评测:非生成式“System One”模型在医学问答上的表现

Jev in Medicine: A Benchmark Evaluation

AI 导读

研究评测了非生成式“System One”模型 Jev 1.13 在 MetaMedQA、PubMedQA、DiagnosisArena-MCQ 和 NEJM Case Challenges 四个医学基准上的表现,以 GPT-6 Sol 为参照。

正文

View PDF HTML (experimental)

Abstract:Jev is a non-generative "System One" model that assigns probabilities to predefined answer options and cannot answer outside them. Its accuracy and calibration on medical question-answering and case-based diagnostic-reasoning tasks are unknown. We evaluated Jev 1.13 on four medical benchmarks: MetaMedQA, PubMedQA, DiagnosisArena-MCQ and the NEJM Case Challenges. GPT-6 Sol, with (medium) and without reasoning, was the reference. The primary outcome was top-1 accuracy; key secondary outcomes were calibration, selective prediction and recognition of unanswerable questions. All 8,469 requests returned a valid answer. Jev's accuracy was similar to that of GPT-6 Sol with medium reasoning on PubMedQA (78.4% vs 78.2%;), lower on MetaMedQA (74.8% vs 82.7%) and much lower on DiagnosisArena-MCQ (59.8% vs 82.4%;) and the NEJM cases (61.8% vs 82.4%). On MetaMedQA, Jev's probabilities were the best calibrated (expected calibration error 0.063 vs 0.146), and its answers with a probability of at least 0.9 (52.9% of questions) were 93.4% accurate, but GPT-6 Sol was as accurate when it accepted a similar proportion of questions. On DiagnosisArena-MCQ, Jev's probabilities discriminated poorly (AUROC 0.645 vs 0.768). Of the 162 questions whose correct answer was "I don't know or cannot answer", Jev chose that option for 10.5% (GPT-6 Sol, 8.6%). Median latency was 0.27-0.31 s; all 2,823 items cost USD 0.08. Jev was fast and inexpensive, and its accuracy was similar to that of a frontier LLM on research abstracts but lower on examination questions and much lower on complex diagnostic cases. Task-specific validation is required before clinical use.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2609.34024 [cs.AI]
  (or arXiv:2609.34024v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.34024

arXiv-issued DOI via DataCite

Submission history

From: Alfredo Madrid Garcia [view email]
[v1] Sun, 27 Sep 2026 23:33:04 UTC (753 KB)
[v2] Thu, 1 Oct 2026 06:03:21 UTC (753 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org