跳到正文
arXiv:cs.CL· Shivani Kumar, Adarsh Bharathwaj, David Jurgens·· 6 小时前AI 评分53

研究:行为博弈测得的合作画像可预测多智能体 LLM 团队表现

Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows

AI 导读

一项被 COLM 2026 接收的研究对 41 个开源权重 LLM 进行六种行为经济学博弈测试,发现博弈得出的合作画像能稳健预测多智能体团队在 AI for Science 任务中的表现。

正文

View PDF HTML (experimental)

Abstract:Multi-agent systems built from teams of large language models (LLMs) are increasingly deployed for collaborative scientific reasoning and problem-solving. These systems require agents to coordinate under shared constraints, such as GPUs or credit balances, where cooperative behavior matters. Behavioral economics provides a rich toolkit of games that isolate distinct cooperation mechanisms, yet it remains unknown whether a model's behavior in these stylized settings predicts its performance in realistic collaborative tasks. Here, we benchmark 41 open-weight LLMs across six behavioral economics games and show that game-derived cooperative profiles robustly predict downstream performance in AI-for-Science tasks, where teams of LLM agents collaboratively analyze data, build models, and produce scientific reports under shared budget constraints. Models that effectively coordinate in games and invest in multiplicative team production (rather than greedy strategies) produce better scientific reports across three outcomes, accuracy, quality, and completeness. These associations hold after controlling for multiple factors, indicating that cooperative disposition is a distinct, measurable property of LLMs not reducible to general ability. Our behavioral games framework thus offers a fast diagnostic for screening cooperative fitness before costly multi-agent deployment.
Comments: Accepted at COLM 2026
Subjects: Computation and Language (cs.CL); Computers and Society (cs.CY); Multiagent Systems (cs.MA)
Cite as: arXiv:2604.20658 [cs.CL]
  (or arXiv:2604.20658v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2604.20658

arXiv-issued DOI via DataCite

Submission history

From: David Jurgens [view email]
[v1] Wed, 22 Apr 2026 15:07:54 UTC (132 KB)
[v2] Tue, 6 Oct 2026 17:29:29 UTC (176 KB)

来源:arXiv:cs.CL · arxiv.org