arXiv:cs.AI· Baihan Lin·· 5 小时前AI 评分58
arXiv 论文:语言模型对抑郁的评分更多反映评分者而非患者
Language-model ratings of depression reflect the rater more than the patient
AI 导读
Baihan Lin 在 arXiv 论文(arXiv:2610.08501)中预注册了 880 个语言模型评分者,将 11 个开源模型与不同提示词和评分方式交叉,应用于 189 段访谈并对照 PHQ-8。
正文
Abstract:Depression has no diagnostic blood test. Language models promise tireless, consistent assessment, but can accurate raters disagree about individuals? We pre-registered 880 language-model raters, crossing 11 open models with prompting and scoring choices, and applied them to 189 interviews against the eight-item Patient Health Questionnaire. Model choice explained 30.0% of summed-symptom score variance, stable participant differences 10.5%. Two randomly drawn raters with area under the receiver operating characteristic curve (AUC) >= 0.70 disagreed on screening decisions for 40% of participants, on average. Average over-rating governed how many were flagged, yet equal-capacity raters chose differently for about one participant in five. A locked analysis of 86 new interviews reproduced the main pre-registered findings. Exploratory recalibration with 40 labelled participants raised accuracy from about 60% to 75% and halved disagreement, leaving one participant in five decided differently. Calibration repaired much of the rater dependence without securing agreement about individuals.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Neurons and Cognition (q-bio.NC) |
| Cite as: | arXiv:2610.08501 [cs.CL] |
| (or arXiv:2610.08501v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08501 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Baihan Lin [view email]
[v1]
Tue, 6 Oct 2026 15:08:17 UTC (823 KB)
来源:arXiv:cs.AI · arxiv.org