arXiv:cs.AI· David Fraile Navarro, Jialei Sheng, Farah Magrabi, Enrico Coiera·· 4 小时前AI 评分61
arXiv 研究:评测格式而非模型能力主导 ChatGPT Health 分诊失败率
Evaluation format, not model capability, drives measured triage failure in the assessment of consumer health AI
AI 导读
arXiv 论文(arXiv:2603.11413)指出,此前 Nature Medicine 报告的 ChatGPT Health 对 51.6% 急症漏分诊主要源于评测格式而非模型能力。
正文
Abstract:A recent Nature Medicine study reported that ChatGPT Health under-triages 51.6% of emergencies and concluded that consumer-facing AI triage poses safety risks. Its protocol, however, was an exam-style scaffold (forced A/B/C/D output, knowledge suppression, no clarifying questions) unlike how consumers use health chatbots. We ask whether the headline error rate is a property of the models or of the measurement. In a first, mechanistic study, five frontier LLMs on a 17-scenario bank scored 6.4 points higher under naturalistic patient-style messages than under the constrained scaffold (p=0.015), and on one vignette three models went from 0-24% with forced choice to 100% with free text. In a second, faithful replication we ran the authors' own 60 released vignettes through six frontier models under four matched formats, with clinician validation of the rewrites and a blinded clinician audit of the LLM adjudicators. Here the direction reversed: free-text rewrites scored slightly below the exact structured prompt (78.7% vs 81.8%, p=0.020), removing only the answer scaffold changed little (80.6% vs 81.4%, p=0.81), and naturalistic input with a forced categorical answer beat both the exact prompt (84.4%, p=0.023) and free text (p=0.0001). On the four vignettes defining the original emergency rate, under-triage was 17% with the exact scaffold, 50% in free text, and 21% when the same message was answered with a forced letter; every free-text "under-triage" was a same-day recommendation scored C rather than D, and 60% made escalation conditional on information the patient was asked to check, behavior a single-turn benchmark cannot score. The headline rate is therefore largely a property of output format and of mapping prose onto a four-point scale. Benchmark scaffolds are behaviorally active instruments; safety claims should report sensitivity to input wording, output format and adjudication.
| Comments: | 10 pages |
| Subjects: | Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2603.11413 [cs.HC] |
| (or arXiv:2603.11413v4 [cs.HC] for this version) | |
| https://doi.org/10.48550/arXiv.2603.11413 arXiv-issued DOI via DataCite |
Submission history
From: David Fraile Navarro MD PhD [view email]
[v1]
Thu, 12 Mar 2026 00:58:22 UTC (12 KB)
[v2]
Sun, 15 Mar 2026 12:43:26 UTC (13 KB)
[v3]
Thu, 26 Mar 2026 02:43:59 UTC (13 KB)
[v4]
Fri, 2 Oct 2026 02:36:04 UTC (24 KB)
来源:arXiv:cs.AI · arxiv.org