跳到正文
arXiv:cs.AI· Zihan Zhou, Xinzhe Hu, Hanxu Yang, Liangjian Wen, Zhao Kang·· 6 小时前AI 评分41

S1-MAS:用 System One 引导的计算分工实现 token 高效的多智能体协作

Token-Efficient Multi-Agent Collaboration via System One-Guided Computational Division of Labor

AI 导读

研究者提出 S1-MAS,一个基于 System One 引导计算分工的 token 高效多智能体框架,将任务选择、角色分配、消息路由等有界协调决策交给轻量 System One 模型,仅把开放式推理留给强 LLM。

正文

View PDF HTML (experimental)

Abstract:Large language model (LLM)-based multi-agent systems (MAS) have become a promising paradigm for complex information-seeking and reasoning tasks by enabling collaborative problem solving among specialized agents. However, existing MAS frameworks tightly couple task reasoning with coordination operations, including task selection, role assignment, message routing, and context management. As interactions grow, using powerful LLMs for these bounded control decisions introduces substantial token overhead and latency, limiting the scalability of agentic Web services. In this paper, we investigate whether coordination can be decoupled from expensive reasoning without compromising collaborative performance. We propose S1-MAS, a token-efficient multi-agent framework based on System One-guided computational division of labor. S1-MAS assigns bounded coordination decisions to lightweight System One models while reserving open-ended reasoning for capable LLM workers. Specifically, a lightweight controller selects inspection conditions, chooses subsequent tasks, and determines termination, while a compact reader retrieves condition-relevant evidence from authorized sources to support these decisions. Through a decision-evidence loop, selected tasks dynamically determine worker roles and source access, enabling adaptive collaboration without task-specific training. Extensive experiments on seven diverse benchmarks demonstrate that S1-MAS achieves superior accuracy while substantially reducing the inference cost. Across individual comparisons with AgentVerse, DyLAN, and SelfOrg on seven benchmarks, S1-MAS reduces GPT-4o token consumption by 44.9%-97.2% and measured end-to-end latency by 37.8%-93.0%. These results highlight its potential for scalable and cost-effective agentic Web applications.
Subjects: Multiagent Systems (cs.MA); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.08155 [cs.MA]
  (or arXiv:2610.08155v1 [cs.MA] for this version)
  https://doi.org/10.48550/arXiv.2610.08155

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhao Kang [view email]
[v1] Tue, 6 Oct 2026 11:11:05 UTC (326 KB)

来源:arXiv:cs.AI · arxiv.org