跳到正文
arXiv:cs.CL· Nicole Summer Hsing, Asuka Yuxi Zheng, Yi Zhao, Haoqin Tu, Jen-tse Huang·· 4 小时前AI 评分53

You Only Align Once:通过种子智能体在多智能体系统中传播对齐行为

You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents

AI 导读

arXiv 论文提出 Alignment Propagation:一个对齐后的种子智能体可仅通过自然语言交互将合作行为传播给未修改的智能体。

正文

View PDF HTML (experimental)

Abstract:Ensuring aligned agent behaviors in distributed open multi-agent systems remains challenging, especially as populations grow and unaligned agents may exist. We show that a single aligned agent can propagate cooperative behaviors to unmodified agents purely through natural-language interaction, a phenomenon we term Alignment Propagation. We study this in the Red-Black Game, a team-based iterated Prisoner's Dilemma in which teammates deliberate and vote to determine their team's collective action. By distilling the cooperative reasoning and persuasive dialogues of a teacher model into Qwen3-14B, we obtain a seed agent that, when placed among four unmodified teammates, more than doubles the cooperation rate from 24.8% to 62.2%, outperforming the teacher model and a vanilla Gemini-3.1-Pro. Remarkably, a seed trained exclusively on the Red-Black Game transfers zero-shot to Sugarscape, a spatially grounded survival simulation with pairwise trading, achieving a 91.5% trade success rate versus a 21.6% baseline. Our results reframe multi-agent alignment from an exhaustive per-agent training problem to a scalable social capability that can be engineered through strategic seed placement.
Subjects: Multiagent Systems (cs.MA); Computation and Language (cs.CL)
Cite as: arXiv:2605.27586 [cs.MA]
  (or arXiv:2605.27586v2 [cs.MA] for this version)
  https://doi.org/10.48550/arXiv.2605.27586

arXiv-issued DOI via DataCite

Submission history

From: Nicole Hsing [view email]
[v1] Tue, 26 May 2026 18:56:02 UTC (1,978 KB)
[v2] Fri, 2 Oct 2026 02:31:40 UTC (1,972 KB)

来源:arXiv:cs.CL · arxiv.org