跳到正文
arXiv:cs.AI· David Fraile Navarro, Berardino Como, Jialei Sheng, Soundariya Ananthan, Shlomo Berkovsky·· 4 小时前AI 评分56

LLM 临床分诊选择题格式效应出现在哪里?研究定位到答案选择阶段

Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect

AI 导读

arXiv 论文 arXiv:2605.29889 研究选择题格式导致的 LLM 临床分诊漏判(under-triage)发生在哪个环节。

正文

View PDF HTML (experimental)

Abstract:LLM evaluations using clinician-authored triage vignettes have reported substantial under-triage under constrained multiple-choice testing. Yet model performance on the same clinical cases can change when responses are generated in free text. We test whether this format effect appears while the case is processed or when clinical information is mapped to the final answer. Using sparse-autoencoder (SAE) features in Gemma 3 4B/12B IT and Qwen3-8B, we find that medical features fire on the shared clinical narrative under both formats but are inactive at the multiple-choice decision token. Emergency-tier information is linearly decodable from vignette representations with ROC-AUC $0.95$--$1.00$ under both formats, with no significant format difference, but is attenuated at the decision token. Natural-language autoencoder verbalization and top-feature characterization associate that token with the multiple-choice scaffold. In a direct linear projection, the identified medical features contribute zero, whereas scaffold-peaking features account for over $91\%$ of unsigned attribution in both Gemma models. Behaviorally, whether multiple choice improves or worsens performance depends on the model. Option-order shuffles rule out simple positional bias, and cases that differ between formats are usually one severity tier apart. Together, these findings place the strongest correlates of the format effect at answer selection while leaving open whether unmeasured clinical representations also differ. Code and data to reproduce experiments are available in the study repository. this https URL
Comments: 9 pages main text, 29 pages total including appendices; 7 figures, 25 tables
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2605.29889 [cs.CL]
  (or arXiv:2605.29889v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2605.29889

arXiv-issued DOI via DataCite

Submission history

From: David Fraile Navarro MD PhD [view email]
[v1] Thu, 28 May 2026 13:14:17 UTC (325 KB)
[v2] Fri, 2 Oct 2026 02:05:20 UTC (347 KB)

来源:arXiv:cs.AI · arxiv.org