跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Ishaan Kelkar, Vikram Kakaria, Nebras Alam, Madhur Panwar, Vasu Sharma, Maheep Chaudhary·· 15 小时前AI 评分46

通用人格向量能否媲美针对性引导来抑制 LLM 谄媚?Gemma 2 27B 与 Qwen 3 32B 实验

Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy

AI 导读

研究在 Gemma 2 27B 和 Qwen 3 32B 上,用 Skeptic、Judge、Devil's Advocate 等通用角色人格向量与专门构建的 CAA 谄媚向量对比,在 PhilPapers 基准上"批判性"角色分别达到 CAA 谄媚 logit 降幅的约 68%(Gemma)和 98%(Qwen)。

正文

View PDF HTML (experimental)

Abstract:Language models are often sycophantic: they agree with a user's stated opinion whether or not it is correct. Prior work has shown that this trait can be controlled by steering a model with a sycophancy persona vector (Chen et al., 2025). Such vectors, however, are extracted from data about sycophancy itself. We ask whether we can instead reuse existing vectors for general roles---Skeptic, Judge, Devil's Advocate---that were extracted without targeting sycophancy at all. On Gemma 2 27B and Qwen 3 32B, we compare these role vectors with a purpose-built Contrastive Activation Addition (CAA) sycophancy vector on a largely held-out, counterbalanced PhilPapers benchmark, using task-specific coefficient tuning on a separate split of sycophancy data. The selected "critical" roles achieve, on average, about 68% (Gemma) and 98% (Qwen) of CAA's reduction in the sycophancy logit. Less agreement does not mean more factual errors on the probes we checked: on 16 true and false factual claims, Qwen steered by the Skeptic or Judge vector still gives the correct answer in all 16 cases, matching the unsteered model and CAA. "Conformist" roles do not reliably produce the opposite effect. Role vectors also have low absolute cosine similarity with the measured CAA direction at the layer we steer; they are geometrically separate interventions, although this does not by itself show that they act through distinct downstream mechanisms. Together, these results show that general persona vectors can help mitigate sycophancy in LLMs, even when extracted without sycophancy-specific labels.
Code: this https URL Results: this https URL
Comments: 11 pages. Spotlight at the 2nd Workshop on Epistemic Intelligence in Machine Learning, ICML 2026. Revised manuscript and abstract
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2605.21006 [cs.AI]
  (or arXiv:2605.21006v4 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2605.21006

arXiv-issued DOI via DataCite

Journal reference: 2nd Workshop on Epistemic Intelligence in Machine Learning, ICML 2026 (Spotlight)

Submission history

From: Ishaan Kelkar [view email]
[v1] Wed, 20 May 2026 10:43:17 UTC (189 KB)
[v2] Sat, 6 Jun 2026 15:25:14 UTC (191 KB)
[v3] Sat, 12 Sep 2026 03:58:24 UTC (192 KB)
[v4] Wed, 30 Sep 2026 23:52:26 UTC (681 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org