arXiv:cs.LG· Hashmath Shaik, Gnaneswar Villuri, Alex Doboli·· 4 小时前AI 评分38
能否预测 LLM 评判的位置翻转?基于残差流激活的线性探针研究
Will the Judge Flip? Predicting Position-Sensitive LLM Judgments from Residual Stream Activations
AI 导读
研究发现,初始裁决前记录的残差流激活可预测 LLM 评判因候选回答顺序产生的翻转。在 JudgeBench 的 534 对样本上,针对三个 Qwen3 评判模型和 Llama-3.1-8B 的线性探针达到 .621-.850 AUROC,比结合置信度、logits、回答长度和初始选择的基线高 .062-.113。
正文
Abstract:The order in which candidate responses are presented can change an LLM judge's verdict. Detecting such a position flip ordinarily requires judging each pair in both orders, which doubles the number of judgments. We investigate whether residual stream activations recorded immediately before the initial verdict can predict a flip. We use nested grouped cross-validation to evaluate regularized linear probes on 534 JudgeBench pairs for three Qwen3 judges and Llama-3.1-8B. The linear probes achieve AUROCs of .621-.850 and outperform a combined baseline that uses verbalized confidence, verdict-label logits, response lengths, and the judge's initial choice by .062-.113 AUROC. Linear probes trained on JudgeBench and then frozen achieve AUROCs of .685-.853 on 1,802 MT-Bench comparisons without MT-Bench fitting or recalibration. These results show that pre-verdict activations support prediction of susceptibility to candidate order and outperform the non-activation predictors evaluated here.
| Comments: | It was submitted to neurips workshop (JUDGE workshop) and received a review score 6.5 combined with one strong accept and one above threshold accept |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07115 [cs.LG] |
| (or arXiv:2610.07115v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07115 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hashmath Shaik [view email]
[v1]
Mon, 5 Oct 2026 16:10:03 UTC (13 KB)
来源:arXiv:cs.LG · arxiv.org