跳到正文
arXiv:cs.CL· William Philipp, Finn Fassbender, Daniel Fister, Thorsten Langer, Martje G. Pauly, Rebecca Herzog, Markus A. Hobert, Theresa Paulus, Alexander Baumann, Chi Wang Ip, Lukas L. Goede, Johanna Reimer, Sebastian L\"ons, Ronald B\"ock, Sebastian Fudickar·· 6 小时前AI 评分44

德国开放问答临床基准 MedQADE:LLM 评分者一致性、评估器偏差与弃答研究

Evaluating Large Language Model Raters for German Open-Response Clinical Questions: A Physician-Annotated Benchmark Study of Agreement, Evaluator Bias, and Abstention

AI 导读

研究团队推出 MedQADE,一个含 3,800 组问答、由 10 名医生标注的德语开放式临床问答基准,并用 9 个 LLM 评估器对答案进行打分。医生在答案正确性上一致性中等偏上(平均 Cohen's kappa 0.612),但题目难度一致性有限;Gemini 3 Flash 的评分最接近医生参考(kappa 0.694 vs 0.709)。

正文

Authors:William Philipp, Finn Fassbender, Daniel Fister, Thorsten Langer, Martje G. Pauly, Rebecca Herzog, Markus A. Hobert, Theresa Paulus, Alexander Baumann, Chi Wang Ip, Lukas L. Goede, Johanna Reimer, Sebastian Löns, Ronald Böck, Sebastian Fudickar

View PDF HTML (experimental)

Abstract:Background: Expert-annotated benchmarks for non-English open-response clinical questions are scarce. LLM-as-a-judge systems may scale evaluation but require validation.
Objective: To introduce MedQADE, a standardized German open-response clinical benchmark with physician reference annotations, and evaluate LLM-as-a-judge alignment, self- and intra-family bias, and abstention.
Methods: The benchmark contains 3,800 question-answer sets with answers from five student LLMs and annotations from 10 physicians. All 10 rated the 200-question core; two primary raters assessed each of 3,600 extension questions, with the tenth resolving disagreements. Nine LLM evaluators assessed all sets. We assessed physician reliability, student-model accuracy, evaluator alignment, bias, and abstention.
Results: Physicians showed moderate-to-substantial agreement on answer correctness (unweighted mean pairwise Cohen's kappa = 0.612) but limited agreement on question difficulty (Krippendorff's alpha = 0.208 using squared numeric-score distances). Student-model accuracy was 17.8%-66.0% and generally decreased with physician-rated difficulty. Gemini 3 Flash approached the leave-one-out physician reference (kappa = 0.694 vs 0.709). Four of five models rated their own responses more favorably than out-of-family evaluators; five of six intra-family comparisons were positive. Physician abstention increased with perceived difficulty. Seven of nine LLM evaluators abstained in no more than 0.51% of evaluations; the two strongest evaluators assigned definitive labels to every response.
Conclusions: Strong LLM evaluators approached physician agreement, but evaluator bias and low observed abstention warrant physician validation and further assessment of selective deferral before fully automated evaluation. These results do not establish clinical safety.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2607.01103 [cs.CL]
  (or arXiv:2607.01103v3 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2607.01103

arXiv-issued DOI via DataCite

Submission history

From: William Philipp [view email]
[v1] Wed, 1 Jul 2026 15:55:31 UTC (2,253 KB)
[v2] Fri, 31 Jul 2026 12:56:37 UTC (2,158 KB)
[v3] Tue, 6 Oct 2026 11:06:26 UTC (555 KB)

来源:arXiv:cs.CL · arxiv.org