跳到正文
arXiv:cs.LG· Bolian Li, Ting-Yao Hu, Cheng-Yu Hsieh, Sanjoy Chowdhury, Oncel Tuzel, Raviteja Vemulapalli·· 4 小时前AI 评分33

为智能体强化学习构建 MoE 专家选择结构

Structuring MoE Expert Selection for Agentic Reinforcement Learning

AI 导读

研究提出一种面向智能体任务的层次化路由控制框架,让专家选择在轮次层面与 READ、UPDATE 等智能体操作对齐,并在 token 层面保持局部一致性,同时引入熵门控机制解决后训练稳定性问题。该方法在所有评测基准上成功率提升超过 10 个百分点,表明智能体轨迹结构可作为 RL 后训练中优化 MoE 容量的有效信号。

正文

View PDF HTML (experimental)

Abstract:Long-horizon LLM agents are frequently implemented using sparse mixture-of-experts (MoE) models, yet the co-design of agentic behavior and MoE structures remains underexplored. In this work, we comprehensively study the connections between agentic post-training and MoE expert selection. In off-the-shelf MoE models, we observe expert selection exhibits a specialized structure that naturally aligns with agentic trajectories. Specifically, expert routing overlaps more between turns where the agent performs semantically similar operations (e.g., READ, UPDATE) than between turns with differing operations. However, standard RL algorithms ignore this specialization, allowing the MoE routing to go uncontrolled during training, which empirically limit task performance and inference efficiency. To address this, we introduce a hierarchical routing control framework for agentic tasks. We explicitly encourage turn-level expert selections to align with agentic operations while regularizing token-level expert selections to maintain local consistency. To resolve stability issues that arise during post-training with the proposed methods, we further introduce an entropy-gated control mechanism. Overall, our routing control framework achieves over 10-point improvements in success rate on all evaluated benchmarks. These results demonstrate that agentic trajectory structure provides an effective signal for optimizing MoE capacity during RL post-training.
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as: arXiv:2610.07332 [cs.LG]
  (or arXiv:2610.07332v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07332

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ting-Yao Hu [view email]
[v1] Mon, 5 Oct 2026 20:06:47 UTC (15,433 KB)

来源:arXiv:cs.LG · arxiv.org