arXiv:cs.AI· Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman·· 6 小时前AI 评分49
JEV-as-a-Judge:JEV 用置信度路由判断,何时接受、何时升级给推理型评审
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
AI 导读
JEV-as-a-Judge 提出用只输出标签概率的决策型评审 JEV 做评测,由其置信度决定直接接受结果,还是升级给推理型评审。在十六个生成式与奖励模型评审的对比中,JEV 在可从文本直接读出判断的场景下与 GPT-6 相差不到 3 分,费用仅为 GPT-6 的 0.36%,中位延迟 0.15 秒,但在数学、代码和逻辑等需推导结论的场景落后。
正文
Abstract:LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly. Confidence routing weakens on style-adversarial pairs and reference-free prose; we close with a simple recipe for validating thresholds locally.
| Comments: | Expanded the dataset, updated the results and figures, and added new analyses. The previous result reporting 99% of GPT performance at 57% of the cost is retained in the appendix |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.26550 [cs.AI] |
| (or arXiv:2609.26550v4 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2609.26550 arXiv-issued DOI via DataCite |
Submission history
From: Yubo Li [view email]
[v1]
Tue, 22 Sep 2026 15:05:56 UTC (2,824 KB)
[v2]
Sun, 27 Sep 2026 19:00:59 UTC (1,802 KB)
[v3]
Tue, 29 Sep 2026 15:32:52 UTC (1,605 KB)
[v4]
Tue, 6 Oct 2026 00:32:50 UTC (1,605 KB)
来源:arXiv:cs.AI · arxiv.org