arXiv:cs.LG· Dani Roytburg, Matthew Bozoukov, Matthew Nguyen, Jou Barzdukas, Simon Fu, Narmeen Oozeer·· 4 小时前AI 评分40
打破镜像:用激活引导向量缓解 LLM 评估器的自我偏好偏见
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
AI 导读
研究提出用轻量级引导向量在推理时缓解 LLM 评估器的自我偏好偏见,无需重新训练,采用 Contrastive Activation Addition(CAA)和基于优化的方法构建,最多可降低 97% 的不合理自我偏好。但其在合理自我偏好与无偏见一致上表现不稳定,暗示自我偏好跨越多个或非线性方向。
正文
Abstract:Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models. This bias undermines fairness and reliability in evaluation pipelines, particularly for tasks like preference tuning and model routing. We investigate whether lightweight steering vectors can mitigate this problem at inference time without retraining. We introduce a curated dataset that distinguishes self-preference bias into justified examples of self-preference and unjustified examples of self-preference, and we construct steering vectors using two methods: Contrastive Activation Addition (CAA) and an optimization-based approach. Our results show that steering vectors can reduce unjustified self-preference bias by up to 97\%, substantially outperforming prompting and direct preference optimization baselines. Yet steering vectors are unstable on legitimate self-preference and unbiased agreement, implying self-preference spans multiple or nonlinear directions. This underscores both their promise and limits as safeguards for LLM-as-judges and motivates more robust interventions.
| Comments: | Presented at {Mechanistic Interpretability, Evaluations, Reliable-ML} Workshops, NeurIPS 2025 |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2509.03647 [cs.CL] |
| (or arXiv:2509.03647v3 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2509.03647 arXiv-issued DOI via DataCite |
Submission history
From: Daniel Roytburg [view email]
[v1]
Wed, 3 Sep 2025 18:52:55 UTC (176 KB)
[v2]
Mon, 22 Jun 2026 18:38:14 UTC (241 KB)
[v3]
Mon, 5 Oct 2026 22:52:55 UTC (241 KB)
来源:arXiv:cs.LG · arxiv.org