跳到正文
arXiv:cs.LG· Gowthamkumar Nandakishore·· 4 小时前AI 评分42

CLM-as-a-Judge:开放对比决策模型 CLM-v0.1-8B 在公开评测基准上的裁判能力评估

CLM-as-a-Judge: Evaluating an Open Contrastive Decision Model on Public Judge Benchmarks

AI 导读

开放对比决策模型 CLM-v0.1-8B 在公开裁判基准上接近随机水平:最佳四选一得 0.351(随机为 0.250),两两比较得 0.593(随机为 0.500),在 RM-Bench 与 JudgeBench 上与抛硬币在统计上无法区分,HaluEval 全部输出同一标签。

正文

View PDF HTML (experimental)

Abstract:An open contrastive decision model is near chance as a judge on the hard public benchmarks: Contrastive-LM/CLM-v0.1-8B scores between 0.351 (best- of-four, chance 0.250) and 0.593 (pairwise, chance 0.500), is statistically indistinguishable from coin flipping on RM-Bench and JudgeBench, and answers every HaluEval item with one constant label, matching the trivial always-first baseline at 0.581. Judges with the same parameter count score far higher everywhere: a reward model reaches 0.764 to 0.976 and a generative judge 0.611 to 0.778, and every gap to CLM is significant after Benjamini-Hochberg correction. Two properties do work. Raw confidences are overconfident by up to +0.401, yet one pooled temperature fit on held-out calibration items repairs expected calibration error to at most 0.062, and the repaired confidence ranks the model's own errors above chance on three of six benchmarks. The decision order-flip rate is 0.0002 against 0.2188 for the generative judge, and the length-preference shift is -0.023 against -0.217. The confidence-gated cascade, however, escalates between 0.923 and 1.000 of items to the strong judge at the preregistered 0.97 retention bar: calibrated confidence about a near-chance judge has almost nothing to keep. The design: five public preference benchmarks and one hallucination benchmark with real labels, scored under a preregistration frozen before any test item was seen, against generative, reward-model, and trivial baselines, with per-item predictions released.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.07177 [cs.LG]
  (or arXiv:2610.07177v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07177

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Gowtham Kumar Nanda Kishore [view email]
[v1] Mon, 5 Oct 2026 18:01:37 UTC (2,955 KB)

来源:arXiv:cs.LG · arxiv.org