跳到正文
arXiv:cs.AI· Baihan Lin·· 5 小时前AI 评分58

arXiv 论文:语言模型对抑郁的评分更多反映评分者而非患者

Language-model ratings of depression reflect the rater more than the patient

AI 导读

Baihan Lin 在 arXiv 论文(arXiv:2610.08501)中预注册了 880 个语言模型评分者,将 11 个开源模型与不同提示词和评分方式交叉,应用于 189 段访谈并对照 PHQ-8。

正文

View PDF HTML (experimental)

Abstract:Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A locked analysis of 86 new interviews reproduced the main pre-registered findings. Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently. Calibration repaired much of the rater dependence without securing agreement about individuals.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Neurons and Cognition (q-bio.NC)
Cite as: arXiv:2610.08501 [cs.CL]
  (or arXiv:2610.08501v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08501

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Baihan Lin [view email]
[v1] Tue, 6 Oct 2026 15:08:17 UTC (823 KB)

来源:arXiv:cs.AI · arxiv.org