arXiv:cs.CL· Longwei Cong, Sonja Hahn, Sebastian Gombert, Leon Camus, Fabian Zehner, Hendrik Drachsler, Ulf Kroehne·· 4 小时前AI 评分41
超越分数对齐评估 LLM-as-a-Judge:残余评判难度的心理测量学分析
Evaluating LLM-as-a-Judge Beyond Score Alignment: A Psychometric Analysis of Residual Judging Difficulty
AI 导读
研究从心理测量学视角提出"残余难度"指标,用于衡量人类与 LLM 评判者在同一评估案例上是否感到同样困难。在 SummEval 上对 17 个开源权重 LLM 评判者的分析显示,潜在摘要质量的中等对齐并不意味着残余难度对齐,且这种错位强烈依赖评估维度:一致性维度上 LLM 更难,连贯性维度上人类更难。人类易而 LLM 难的案例可部分由源文本与摘要的可观测属性预测。
正文
Abstract:Large language models (LLMs) are widely used as automatic judges, with validity typically assessed via alignment with human scores. However, aggregate agreement fails to reveal whether humans and LLMs find the same evaluation cases difficult. In this paper, we study this problem in summarization evaluation from a psychometric perspective. We fit Many-Facet Rasch Models separately to human and LLM ratings to decompose scores into latent summary quality, rater severity, dimension severity, and rating-scale thresholds. Building on this decomposition, we define residual hardness as a model-adjusted measure of judging difficulty and compare whether human and LLM judges share the same hardness structure. Across 17 open-weight LLM judges on SummEval, we find that moderate alignment in latent summary quality does not imply alignment in residual hardness. Human and LLM judges differ in which summary--dimension units remain difficult, and this mismatch is strongly dimension-dependent. Consistency shows a pronounced LLM-hard shift, whereas coherence shows a human-hard shift. We further show that human-easy but LLM-hard cases are partially predictable from observable source--summary properties. These findings suggest that aggregate human alignment reflects only part of LLM-as-a-judge reliability, while psychometric residual diagnostics support more informative judge evaluation and more targeted human--LLM collaboration.
| Comments: | Accepted at AACL-IJCNLP 2026 |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.02877 [cs.CL] |
| (or arXiv:2610.02877v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02877 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Longwei Cong [view email]
[v1]
Fri, 2 Oct 2026 06:13:00 UTC (4,092 KB)
来源:arXiv:cs.CL · arxiv.org