跳到正文
原文
elvis· @omarsar0 · X·· 6 小时前AI 评分46
AI 导读

AgentWorld 将 3 到 20 个不同角色的 LLM 智能体放入游戏沙盒,执行 50+ 轮的长程任务,智能体无法看到彼此内部状态,只能通过消息和共享计划协调。

正文

More agents don't mean higher performance.

There is a coordination bottleneck to consider. Not to mention the unnecessary costs.

So how many of a multi-agent team's actions actually help it finish the task?

In this AgentWorld paper, fewer than a third.

AgentWorld puts 3 to 20 LLM agents with different roles into a game sandbox for tasks that run 50+ rounds.

Agents can't see each other's internal state, so they have to coordinate through messages and shared plans.

Gemini 3 Flash has the highest task success at 52.0%. Coordination tasks are the hardest category, at 12% success, and common failures include communication breakdowns, role confusion, and lost shared plans.

Paper: https://arxiv.org/abs/2609.31590

Chat with Paper: https://academy.dair.ai/papers/agentworld-benchmarking-long-horizon-collaboration-of-multi-agent-llms-2609.31590

来源:elvis · x.com