arXiv:cs.CL· Vineet Kumar, Darshita Rathore, Anindya Moitra·· 3 小时前
所有评判并非等价:重新思考 LLM Judge 的可靠性
All Verdicts are Not Equal: Rethinking LLM Judge Reliability
AI 导读
一项针对 LLM-as-a-Judge 的可靠性审计发现,六个前沿模型在温度为零的重复评测中仍会出现判定变化,位置顺序互换会翻转多数困难任务的判定结果。研究提出可信判定率(T)指标,用于衡量评测可复现、顺序不变且准确的联合概率,并指出可靠性是项目特异而非模型级别的。从成对胜率转向整体评分标准比任何单一提示词干预都更能提升可信度。
正文
Abstract:LLM-as-a-Judge is the standard paradigm for NLP evaluation, yet its systemic reliability remains poorly understood despite being widely treated as a deterministic ground truth. We present a comprehensive reliability audit, stresstesting six frontier models across four benchmarks, five prompt formats, two presentation orders, three sampling temperatures, and ten repetitions per condition. Our empirical analysis reveals severe vulnerabilities: verdicts change across identical replications at temperature zero, position-order swaps flip the majority of verdicts on challenging tasks, and the most deterministic judge achieves perfect consistency by trivially repeating incorrect verdicts, agreeing with ground truth only 51% of the time. To formalize these multi-faceted failure modes, we introduce the trustworthy verdict rate (T ), a unified metric capturing the joint probability that an evaluation is reproducible, order-invariant, and accurate. UsingT , we derive a theoretical upper bound on accuracy imposed by position bias and show that reliability is item-specific rather than modellevel. Finally, we demonstrate that shifting from pairwise win-rate to holistic rubric scoring improves trustworthiness more than any single-format prompting intervention, offering an actionable framework for robust NLP evaluation.
| Comments: | Accepted at AACL IJCNLP (Main) 2026 |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.12083 [cs.CL] |
| (or arXiv:2610.12083v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12083 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Vineet Kumar [view email]
[v1]
Thu, 8 Oct 2026 14:55:29 UTC (732 KB)
来源:arXiv:cs.CL · arxiv.org