跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang·· 15 小时前AI 评分41

开源大语言模型临床分诊偏见反事实审计:Qwen2.5、MedGemma、GPT-OSS 等十款模型对比

Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage

AI 导读

一项针对十款开源大语言模型的反事实审计显示,模型对人口统计学、社会经济等变量的敏感性并未随参数规模增大或医学领域预训练而一致降低。微调后的 Qwen2.5-7B 整体敏感性最低,任意偏移率为 5.27%、平均绝对偏移 0.0534,而基座模型分别为 16.02% 和 0.1706;部分更大或医学领域模型偏移更显著。该框架可用于临床部署前比较开源 LLM 的公平性风险。

正文

View PDF HTML (experimental)

Abstract:Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) are increasingly considered for local and privacy-preserving clinical decision support, it remains unclear how counterfactual bias varies across model families, sizes, medical-domain models, and domain-adapted models. We present a comparative counterfactual audit of ten open-source LLMs for pediatric Emergency Severity Index (ESI) prediction. Starting from real and handbook-style clinical vignettes, we construct paired counterfactual variants that change only one injected demographic, socioeconomic, healthcare-access, behavioral, social, or system-context variable while holding the clinical presentation fixed. Models include Qwen2.5-7B, Qwen2.5-14B-Instruct, a QLoRA fine-tuned Qwen2.5-7B, MedGemma variants, MedLLaMA2-7B, GPT-OSS-20B, and GPT-OSS-120B. We measure any counterfactual shift, undertriage, overtriage, shifts greater than one ESI level, mean shift, and mean absolute shift. Counterfactual sensitivity varied substantially and did not consistently decrease with larger model size or medical-domain pretraining. The fine-tuned Qwen2.5-7B showed the lowest overall sensitivity, with a 5.27% any-shift rate and mean absolute shift of 0.0534, versus 16.02% and 0.1706 for the base model. Several larger or medical-domain models showed more significant shifts. Stratified and correlation analyses further revealed clinically important directionality and shared failure patterns hidden by aggregate rates. These findings support counterfactual auditing as a lightweight, clinically interpretable framework for comparing fairness risks in open-source LLMs before clinical deployment.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.01963 [cs.AI]
  (or arXiv:2610.01963v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.01963

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Manar Aljohani [view email]
[v1] Thu, 1 Oct 2026 16:17:43 UTC (2,093 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org