跳到正文
arXiv:cs.LG· Nicolas Martorell, Wendy Brau, Gonzalo A. Heredia, Tom\'as Pablo Korenblit, Gaspar Labasti\'e, Tom\'as Gimenez Molina·· 3 小时前AI 评分44

PowerBench:衡量语言模型在权力转移请求中的偏见

PowerBench: Measuring Language Model Bias in Power-shifting Requests

AI 导读

研究者推出 PowerBench,用于评估语言模型在权力转移请求中的偏见,区分自我赋权、去权和夺权三类请求,并构建了开源数据集,覆盖不同权力领域、情境、受影响方规模及用户原有权力地位。

正文

View PDF HTML (experimental)

Abstract:Language models increasingly assist people with power-related requests, so systematic differences in whom they help could shift the distribution of power at scale, or be exploited by users who learn which identities are refused less. We introduce PowerBench, an evaluation of power-shifting requests that distinguishes self-empowerment, disempowerment, and power grabbing, plus a control of refusal-inducing requests that shift no power. We build, curate, and open-source a dataset of such requests varying the power domain, the context, the scale of the affected party, and the prior power standing of the user, and evaluate 24 models (12 from US and 12 from Chinese developers) under three experimental conditions: reciprocal nationalities of user and affected party, an AI agent as the user, and 8 request languages. Models refuse power grabbing more than disempowerment, and disempowerment more than self-empowerment. Refusal of power grabbing rises with the scale of the affected party, from an individual to a society. Models are biased toward helping others take power from the US and against helping US users take power from others, but favor the US when it gains power and nobody loses it. When the user is an AI agent, refusal of power-shifting requests increases, especially in power grabbing against an individual. Finally, language biases refusal, but in model-specific ways that largely cancel on average. We release PowerBench to make these asymmetries measurable in current and future models.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02303 [cs.LG]
  (or arXiv:2610.02303v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02303

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Nicolas Martorell [view email]
[v1] Thu, 1 Oct 2026 17:57:25 UTC (840 KB)

来源:arXiv:cs.LG · arxiv.org