跳到正文
原文
arXiv:cs.AI(全量分类)· Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen·· 5 小时前AI 评分53

多用户多智能体团队为何表现更差:arXiv 论文提出 MAMUBench 评测基准

Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

AI 导读

arXiv 论文(arXiv:2610.00583)研究多用户多智能体协作,在五个前沿模型、四个环境共 77 个场景中比较单一协调者智能体与各服务一名用户的智能体团队。

正文

View PDF HTML (experimental)

Abstract:People are increasingly delegating tasks to AI agents, and those agents are increasingly encountering other people's agents over shared resources such as a codebase, a calendar, or a budget. When each agent acts for a different user with different goals, coordination often fails, and the group ends up worse off than if a single agent had acted for everyone. We study this multi-user, multi-agent setting across five frontier models and 77 scenarios in four environments: an API key environment in which agents share a compute budget, a clinic in which they share a calendar, a personal assistant environment in which they share a group order or booking, and a merge queue in which they share a release cutoff. In each scenario, we compare a single agent that serves every user (a coordinator) to a team in which each agent serves one user, with and without a communication channel between the agents. Teams deliver worse group outcomes than the coordinator in every environment: without a channel, they completely collapse in two environments, and even with one, coordination overhead creates substantial gaps. For example, in the personal assistant environment, the coordinator fulfills a targeted user request about twice as often as teams. We identify distinct behaviors associated with this poor group-level performance, including stalling as teams grow, overriding each other's actions, and fabricating claims. We find effective but environment-specific mitigations, such as a team lead, explicit procedural instructions, and a platform check that makes an agent read its peers' messages before committing. We will release the API key, clinic, and personal assistant environments as MAMUBench, comprising 74 scenarios for evaluating multi-user, multi-agent coordination.
Comments: 63 pages, 22 Figures, 10 Tables, Code: this https URL (will be released after review)
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.00583 [cs.AI]
  (or arXiv:2610.00583v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.00583

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sahan Paliskara [view email]
[v1] Wed, 30 Sep 2026 18:49:25 UTC (3,747 KB)

来源:arXiv:cs.AI(全量分类) · arxiv.org