跳到正文
arXiv:cs.AI· Mohamed Ayman Mohamed, Harshil Kotamreddy, Marcos Menon Jose·· 6 小时前AI 评分38

自参照社会偏好:无需观察他人奖励即可实现合作

Self-Referenced Social Preferences: Cooperation without Observing Others Rewards

AI 导读

研究提出"自参照社会偏好"方法,让每个智能体学习自身奖励模型,并将其应用于其他智能体观测到的状态转移,从自身视角评估他人结果,无需访问对方的私有奖励信号。在 Escape Room、Clean Up 和 Commons Harvest 三个序贯社会困境中,智能体均学会了合作行为,且比能访问真实奖励的智能体更常实现更公平的收益分配。

正文

View PDF HTML (experimental)

Abstract:Social preferences can promote cooperation in multi-agent reinforcement learning, but existing approaches often require agents to observe the rewards of their peers. In many real-world interactions, however, an agent can, as humans do, observe others' behavior and outcomes without access to their private reward signals. We introduce self-referenced social preferences, in which each agent learns a model of its own reward, applies it to other agents' observed transitions to assess their outcomes from its own perspective, and feeds these self-referenced assessments into standard social preferences. We study two ways to incorporate these assessments: modifying the learning reward, or using them to weight policy updates. We evaluate the approach on three sequential social dilemmas, Escape Room, Clean Up, and Commons Harvest, which require volunteering, public-good contribution, and resource restraint, respectively. Across all three environments, agents learn cooperative behavior without observing others' rewards, including in settings where independent learners fail to cooperate, and frequently achieve more equitable divisions of jointly produced returns than agents with access to true rewards. The effective integration point depends on the social preference: inequity aversion works best in the reward together with a value look-ahead, whereas a purely benevolent preference benefits from policy-update weighting. Under partial observability, the policy-update approach continues to support cooperation. These results show that explicit access to other agents' reward signals is not necessary for learning cooperative behavior: social preferences can instead be grounded in self-referenced assessments of others' outcomes derived from their observed behavior.
Subjects: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA)
Cite as: arXiv:2610.07881 [cs.AI]
  (or arXiv:2610.07881v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07881

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mohamed Mohamed A. [view email]
[v1] Tue, 6 Oct 2026 07:29:38 UTC (1,218 KB)

来源:arXiv:cs.AI · arxiv.org