arXiv:cs.LG· Jabin Koo, Soheil Abbasloo, Sungjae Lee, Jungseul Ok·· 3 小时前AI 评分34
面向多智能体系统的动态专家剪枝(DEP)
Dynamic Expert Pruning for Multi-Agent Systems
AI 导读
研究者提出动态专家剪枝(DEP),利用智能体的系统提示词与任务提示词直接预测每次请求所需的专家,单次前向即可生成专属掩码,无需逐配置校准。该方法在多种任务、角色、模型规模和 MoE 架构上整体准确率优于静态剪枝与合并基线,并能泛化到训练中未见的工作流;保留专家越少时优势越明显,表明多智能体系统的角色特化支持比静态剪枝更稀疏的部署。
正文
Abstract:Mixture-of-Experts (MoE) architectures scale language models efficiently by activating only a few experts per token, but the saving is confined to computation: every expert must stay resident on the accelerator, so memory bounds where these models can be deployed. Expert pruning reduces this footprint, yet existing methods are static --- a single mask, calibrated offline, is applied to the model for every subsequent request. This assumption can fail when the workload is heterogeneous, most prominently in multi-agent systems, where one backbone serves many tasks and roles at once: our analysis shows that different tasks and roles recruit different experts, while static methods assign one fixed subset to all of them. We therefore propose Dynamic Expert Pruning (DEP), which rests on a finding we establish here: an agent's system and task prompts are by themselves sufficient to identify the experts that agent and its task require, since that text already describes what the agent will do. A lightweight predictor, trained once on workflow transcripts, turns those prompts into a specialized per-request mask in a single forward pass, with no per-configuration calibration. Across diverse tasks and roles, model scales, and MoE architectures, DEP achieves better overall accuracy than static pruning and merging baselines, and generalizes to workflows unseen in training without retraining. Its margin over those baselines is largest when few experts are retained, suggesting that the role specialization inherent to multi-agent systems permits sparser serving than static pruning allows.
| Comments: | 18 pages, 3 figures, 15 tables |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA) |
| Cite as: | arXiv:2610.02951 [cs.LG] |
| (or arXiv:2610.02951v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02951 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jabin Koo [view email]
[v1]
Fri, 2 Oct 2026 07:45:30 UTC (375 KB)
来源:arXiv:cs.LG · arxiv.org