arXiv:cs.AI· Yujiao Chen·· 3 小时前
研究:LLM 智能体为何偏爱自身群体——利害关系、观察到的规范与声誉
Why LLM Agents Favor Their Group: Stakes, Observed Norms, and Reputation
AI 导读
研究在约4400个微型社会、330万次模型调用中测试15个OpenAI模型和3个Claude模型,发现群体标签带来的偏袒只在给他人打分不损害自身利益时出现,一旦有成本便消失。有代价时,互动历史成为偏袒主因,在15个模型中13个呈统计显著正效应,其中11个达10分中的约3.5-8分,并随轮次增长、延伸至陌生同标签者。脚本化历史显示,平等规范下偏袒趋近于零,个体声誉可覆盖群体偏好。
正文
Abstract:Language-model agents favor their own group because they have watched their members favor each other. The group label alone does little once the decision has a cost; what drives favoritism is observed behavior, and an individual's own record can override it. We test this in small societies with arbitrary group labels, ten rounds of point sharing, and matched one-shot decisions across fifteen OpenAI models and three Claude models, about 4,400 societies and 3.3 million audited model calls. First, the large effect of a bare group label reported in earlier work appears only when giving others points costs the agent nothing; once the agent can keep points for itself, that effect collapses on every model that shows it. Second, under a stake, interaction history becomes the main source of favoritism: the history effect is statistically positive on 13 of 15 models, reaches about 3.5-8 points out of 10 on 11, grows with the number of rounds played, and extends to labeled strangers the agent has never met. Third, with scripted histories, favoritism falls to near zero under an egalitarian norm and reverses when the agent's own group is seen favoring the other side; stronger models side with an individual's record when it conflicts with the group. Group favoritism is thus conformity to observed group behavior, carried to strangers by the label and overridden by individual reputation. The same account predicts responses to betrayal, scandal, and a free offer to change group: public reprimand repairs betrayal better than apology or restitution, allocation punishment remains confined to the offending member, and a formed group cannot be bought but can, on weaker models, be invited away.
| Subjects: | Physics and Society (physics.soc-ph); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11008 [physics.soc-ph] |
| (or arXiv:2610.11008v1 [physics.soc-ph] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11008 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yujiao Chen [view email]
[v1]
Wed, 7 Oct 2026 23:46:37 UTC (61 KB)
来源:arXiv:cs.AI · arxiv.org