跳到正文
arXiv:cs.CL· Yiqi Liu, Joseph James, Yang Wang, Kun Zhao, Chenghao Xiao, Chenghua Lin·· 6 小时前AI 评分38

JudgeMoE:面向 LLM-as-a-Judge 的分布聚合方法

JudgeMoE: Distributional Aggregation for LLM-as-a-Judge

AI 导读

JudgeMoE 是一种轻量聚合器,对缓存的 judge 分数分布分配样本级权重并融合后再计算最终分数。

正文

View PDF HTML (experimental)

Abstract:When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression. We introduce JudgeMoE, a lightweight aggregator that assigns example-specific weights to cached judge score distributions and fuses them before computing a final score. A protocol study shows that score-range choice is unstable across judge--dataset settings and that soft scoring usually outperforms hard decoding. On the original 10-cell benchmark, JudgeMoE improves mean Spearman over uniform log pooling by $+0.079$. Applying the same configuration to six additional cells yields a $+0.0393$ mean gain over the strongest local single judge across 16 cells, with positive differences in 12/16 cells and a one-sided Wilcoxon signed-rank $p=0.0091$. Validation-based analyses further show that the preferred aggregation method depends on the task and judge pool.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.07109 [cs.CL]
  (or arXiv:2610.07109v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.07109

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yiqi Liu [view email]
[v1] Mon, 5 Oct 2026 15:41:46 UTC (298 KB)

来源:arXiv:cs.CL · arxiv.org