跳到正文
arXiv:cs.CL· Juan Manuel Contreras·· 4 小时前AI 评分56

LLM 原生心理测量工具揭示 25 个模型的自陈报告与行为差距

An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models

AI 导读

作者构建基于 LLM 专属行为(如过度拒答、主动免责声明)的自陈量表,对来自 17 家开发商的 25 个 LLM 各施测 300 题 30 次,得到五个可复现因子。

正文

View PDF HTML (experimental)

Abstract:Do large language models' (LLMs') answers to self-report questionnaires predict how they behave? Prior work finds they do not, but it uses human personality inventories, so the gap could reflect borrowed human constructs rather than LLM self-report itself. We test this with a self-report instrument built from LLM-specific behaviors (e.g., over-refusal, unsolicited disclaimers) whose structure is derived bottom-up. Administering 300 items 30 times to 25 LLMs from 17 developers yields five replicable, reliable factors (Tucker $\phi \geq .957$, $\alpha \geq .930$). We compare these self-reports with 2,500 open-ended behavioral samples rated by 151 humans and an LLM-judge ensemble. Humans and judges agree about model behavior ($\bar{r} = .51$), but self-report barely tracks human ratings ($\bar{r} = .09$, 95% CI $[-.07, .18]$) or rater-free text measures, and correcting for criterion unreliability leaves four of five factors near zero. Verbosity is the partial exception ($r = .40$, 71% of its reliability ceiling). On Responsiveness, self-report tracks LLM judges more than humans ($r = .53$ vs. $.18$; Steiger $p = .04$), and controlling for length and formatting does not remove this: agreement between LLM judges and LLM self-report is weak evidence that either tracks human judgment.
Subjects: Human-Computer Interaction (cs.HC); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
ACM classes: I.2.7; J.4
Cite as: arXiv:2606.09843 [cs.HC]
  (or arXiv:2606.09843v4 [cs.HC] for this version)
  https://doi.org/10.48550/arXiv.2606.09843

arXiv-issued DOI via DataCite

Submission history

From: Juan Manuel Contreras Ph.D. [view email]
[v1] Fri, 24 Apr 2026 04:42:09 UTC (577 KB)
[v2] Thu, 25 Jun 2026 15:23:14 UTC (571 KB)
[v3] Tue, 7 Jul 2026 00:32:51 UTC (579 KB)
[v4] Wed, 7 Oct 2026 02:52:07 UTC (598 KB)

来源:arXiv:cs.CL · arxiv.org