跳到正文
arXiv:cs.CL· Guang Yang, Homa Hosseinmardi, Fengchen Liu, Amir Ghasemian·· 3 小时前AI 评分59

arXiv 论文:任务、模型与施压方式如何影响大语言模型的谄媚让步

Beyond the Sycophancy Score: How Task, Model, and Pressure Shape LLM Yielding

AI 导读

arXiv 论文分析 LLM 谄媚行为的发生条件,基于 103,939 条标注回复,覆盖 8 个关闭推理的模型和其中 2 个开启最大推理的配置,统一使用 200 个条目、13 种施压条件和四轮对话。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) often abandon a correct answer, or endorse a user's position, once the user pushes back. This behavior, called sycophancy, is usually reported as a single rate per model, which says little about when it happens or how a user can avoid it. We study the conditions that produce it with 103,939 graded replies from ten configurations: eight LLMs with reasoning disabled, and two of them again with maximum reasoning, all facing the same 200 items, 13 pressure conditions, and four-turn conversations, with every reply labeled by two independent LLM judges. We find that the dominant factors are how costly it is for the model to verify the user's claim, and whether a trained guardrail covers it. Removing this task factor from a logistic model costs 0.485 of McFadden $R^2$, against 0.139 for model family and 0.009 for pressure tactic. Anchored facts are almost never conceded (1.3%), while adoption on logic puzzles rises with the number of clues needed to refute the pushed answer. Personal choices are endorsed in 77.0% of conversations. Most concessions on hard items come from models that cannot reliably solve them; models that can solve them rarely give the answer up. For both models tested, maximum reasoning removes these concessions completely: adoption on deep puzzles falls from 19.2% and 12.5% to 0%. Fallacious or emotional framing adds nothing beyond plain repetition. Three human annotators agree with the judges' consensus on 118/120 calibration items. These results give practical rules for reliable use: simplify hard-to-verify problems and reason deeply, state the question rather than one's preferred answer, ask for evidence on open questions, and choose models by their measured guardrail profile.
Comments: Preprint. 27 pages
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.08840 [cs.CL]
  (or arXiv:2610.08840v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08840

arXiv-issued DOI via DataCite

Submission history

From: Guang Yang [view email]
[v1] Wed, 30 Sep 2026 05:27:34 UTC (360 KB)

来源:arXiv:cs.CL · arxiv.org