跳到正文
arXiv:cs.AI· David Fraile Navarro, Jialei Sheng, Farah Magrabi, Enrico Coiera·· 4 小时前AI 评分61

arXiv 研究:评测格式而非模型能力主导 ChatGPT Health 分诊失败率

Evaluation format, not model capability, drives measured triage failure in the assessment of consumer health AI

AI 导读

arXiv 论文(arXiv:2603.11413)指出,此前 Nature Medicine 报告的 ChatGPT Health 对 51.6% 急症漏分诊主要源于评测格式而非模型能力。

正文

View PDF HTML (experimental)

Abstract:A recent Nature Medicine study reported that ChatGPT Health under-triages 51.6% of emergencies and concluded that consumer-facing AI triage poses safety risks. Its protocol, however, was an exam-style scaffold (forced A/B/C/D output, knowledge suppression, no clarifying questions) unlike how consumers use health chatbots. We ask whether the headline error rate is a property of the models or of the measurement. In a first, mechanistic study, five frontier LLMs on a 17-scenario bank scored 6.4 points higher under naturalistic patient-style messages than under the constrained scaffold (p=0.015), and on one vignette three models went from 0-24% with forced choice to 100% with free text. In a second, faithful replication we ran the authors' own 60 released vignettes through six frontier models under four matched formats, with clinician validation of the rewrites and a blinded clinician audit of the LLM adjudicators. Here the direction reversed: free-text rewrites scored slightly below the exact structured prompt (78.7% vs 81.8%, p=0.020), removing only the answer scaffold changed little (80.6% vs 81.4%, p=0.81), and naturalistic input with a forced categorical answer beat both the exact prompt (84.4%, p=0.023) and free text (p=0.0001). On the four vignettes defining the original emergency rate, under-triage was 17% with the exact scaffold, 50% in free text, and 21% when the same message was answered with a forced letter; every free-text "under-triage" was a same-day recommendation scored C rather than D, and 60% made escalation conditional on information the patient was asked to check, behavior a single-turn benchmark cannot score. The headline rate is therefore largely a property of output format and of mapping prose onto a four-point scale. Benchmark scaffolds are behaviorally active instruments; safety claims should report sensitivity to input wording, output format and adjudication.
Comments: 10 pages
Subjects: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI)
Cite as: arXiv:2603.11413 [cs.HC]
  (or arXiv:2603.11413v4 [cs.HC] for this version)
  https://doi.org/10.48550/arXiv.2603.11413

arXiv-issued DOI via DataCite

Submission history

From: David Fraile Navarro MD PhD [view email]
[v1] Thu, 12 Mar 2026 00:58:22 UTC (12 KB)
[v2] Sun, 15 Mar 2026 12:43:26 UTC (13 KB)
[v3] Thu, 26 Mar 2026 02:43:59 UTC (13 KB)
[v4] Fri, 2 Oct 2026 02:36:04 UTC (24 KB)

来源:arXiv:cs.AI · arxiv.org