arXiv:cs.CL· Kamile Dementaviciute, Julija Vaitonyte, Tijl De Bie·· 3 小时前AI 评分46
LLM 说服力评估为何难以一致:九种自动化方法在同十五个 LLM 上排名仅弱相关
LLM Persuasion Is in the Eye of the Evaluation
AI 导读
研究将九种已发表的自动化说服评估方法套用到同一套设置,在同样的十五个 LLM 上运行,发现各方法排名仅弱相关(平均 Spearman ρ = 0.25)。约四分之一的排名分歧来自模型对部分任务拒答,且拒答集中在操纵类任务;多数非操纵的理性说服方法会跟随模型的通用能力,而多数操纵方法不会。
正文
Abstract:Large language models (LLMs) have already been shown to match or exceed human experts in persuasion. While their persuasive capabilities hold promise for beneficial uses such as education and health communication, they can also be used to manipulate and misinform, making their evaluation a growing priority for developers and regulators. That evaluation, however, remains fragmented: studies differ in what they treat as persuasion, and broad claims often rest on narrow, situation-specific assessments. Automated methods, often modelled on human studies, offer a way to compare such assessments directly, as they can be run on the same models at scale and can include high-risk forms of persuasion that would be difficult or unethical to test on people. In this study, we adapt nine published automated methods to a shared setup, run them on the same fifteen LLMs, and ask whether their rankings agree and why. We find that the methods agree only weakly (mean Spearman $\rho = 0.25$). Our analyses point to two contributing factors. Models that refuse some tasks but not others, directly or indirectly, lower agreement by about a quarter, and these refusals fall mostly on manipulation tasks. General capability also plays a part: most rational persuasion (non-manipulative) methods track it, whereas most manipulation methods do not. Together, these findings suggest that agreement depends more on the task a method sets than on how it scores persuasion, although this pattern is only indicative given the eight methods available for analysis. More broadly, our results suggest that persuasion scores combine a model's ability to persuade with its willingness to do so. A single score is therefore informative about its own setting, but says little about a model's persuasiveness across tasks.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY) |
| Cite as: | arXiv:2610.10232 [cs.CL] |
| (or arXiv:2610.10232v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10232 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kamile Dementaviciute [view email]
[v1]
Wed, 7 Oct 2026 15:19:54 UTC (87 KB)
来源:arXiv:cs.CL · arxiv.org