跳到正文
arXiv:cs.AI· Laur\`ene Vaugrante, Thilo Hagendorff·· 5 小时前AI 评分49

LLM-as-a-Judge 的设计选择影响有多大?提示词、评分量表与模型的系统对比

How Much Do LLM-as-a-Judge Design Choices Matter? A Systematic Comparison of Prompt Designs, Rating Scales, and Models

AI 导读

一项研究系统评估了 10 个推理模型在 LLM-as-a-Judge 中的提示词、评分量表和模型选择的影响:评分任务中多数评委可靠(1-7 分制下平均绝对偏差 0.11 分),准确率分类平均达 96.5%。但设计选择会带来偏移,仅更换评分量表即可使测得偏差变化最多 0.93 分,使用详细提示词使宽松度下降 28.9 个百分点,切换模型最多下降 56.1 个百分点,模型身份是方差的主要来源。

正文

View PDF HTML (experimental)

Abstract:Researchers increasingly use Large Language Models as judges (LLM-as-a-judge) to evaluate model outputs. Yet there are no standards for how to design these judges. Typically, researchers choose the prompt, rating scale, and model intuitively. If these choices change the judge's verdicts, two studies can reach different conclusions about the same facts. To address this risk and to provide an empirical basis for judge designs, we evaluate 10 reasoning models across multiple designs on two tasks: a scalar rating of sentence sentiment and toxicity (over 500 items per category), as well as a binary accuracy classification of question-answer pairs (n=600). For the rating tasks, despite judges showing significant disagreements with the human ground truth, the practical size of differences is small enough to consider most judges reliable (mean absolute deviation of 0.11 points on a 1 - 7 scale); toxicity judges even outperform standard classifiers. Judges are also highly accurate on average (96.5%) for the accuracy classification task. However, design choices can produce shifts: changing the rating scale alone can shift measured bias by up to 0.93 points (rating task), and while accuracy levels are rarely impacted, design choices consistently impact judge leniency (classification task; leniency drop of 28.9 percentage points when using detailed prompts, and up to 56.1 percentage points when switching models). Counterintuitively, lower reasoning effort affects neither accuracy nor leniency. Across both tasks, model identity is the dominant source of variance. These findings suggest that while LLM judges are broadly trustworthy in aggregate, design choices can be meaningful sources of variance. Given the growing reliance on automated evaluation in LLM research, we intend this study as a methodological reference for designing more robust and replicable LLM-as-a-judge pipelines.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.05094 [cs.CL]
  (or arXiv:2610.05094v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.05094

arXiv-issued DOI via DataCite

Submission history

From: Laurène Vaugrante [view email]
[v1] Sun, 4 Oct 2026 09:56:48 UTC (592 KB)

来源:arXiv:cs.AI · arxiv.org