arXiv:cs.LG· Tianxi Li, Jie Ding·· 4 小时前AI 评分42
用 AI 评委做可信方法比较:次序、批次与聚合效应下的估计与设计
Trustworthy Method Comparison with AI Judges: Estimation and Design under Order, Batch, and Aggregation Effects
AI 导读
研究发现 LLM 作为评委的评测机制可用一类马尔可夫广义线性混合模型(GLMM)近似,并在三个主流商业 LLM 的样本外预测中得到支持。基于一阶马尔可夫 GLMM:随机化并取平均的排行榜选择在温和分离条件下具有一致性,Williams 方阵设计可在项目质量接近时提升效率;而朴素平均因响应模型的非线性,可能对组间质量差异得出不一致结论。
正文
Abstract:Large language models (LLMs) are increasingly used as judges for automated AI evaluation. A common practice is to randomize prompt sequences and average the resulting scores, but its statistical validity remains unclear. We show that LLM evaluation mechanisms can be approximated by a class of Markov generalized linear mixed models (GLMMs), supported by out-of-sample predictions across three major commercial LLMs. Using a first-order Markov GLMM, we study leaderboard ranking and group comparison. For leaderboard ranking, randomize-and-average selection is consistent under a mild separation condition, and a Williams square design can improve efficiency when item qualities are close. For group comparison, naive averaging can yield inconsistent conclusions about differences in group-level quality because of the response model's nonlinearity. Empirical results further support the validity of the proposed model-based inference beyond the first-order theory, including settings with higher-order sequence memory. We illustrate the approach in an application where AI judges compare two graphical model estimation methods.
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG); Applications (stat.AP); Methodology (stat.ME) |
| Cite as: | arXiv:2610.07755 [stat.ML] |
| (or arXiv:2610.07755v1 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07755 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Tianxi Li [view email]
[v1]
Tue, 6 Oct 2026 04:53:14 UTC (179 KB)
来源:arXiv:cs.LG · arxiv.org