跳到正文
arXiv:cs.AI· Jianglin Qiao, Siyi Hu, Thien Hoang Nguyen, Zehong Cao, Salah Sukkarieh·· 10 小时前AI 评分34

错峰参与下的多智能体强化学习:SPL 训练方法

Cooperating with Future Collaborators: Multi-Agent RL under Staggered Participation

AI 导读

研究提出 Staggered Participation Learning(SPL),针对多智能体强化学习中智能体错峰参与(早期智能体留下对未来决策有用的信息)的场景,用前瞻性获取监督训练早期智能体、用结果监督训练后期智能体。

正文

View PDF HTML (experimental)

Abstract:In cooperative Multi-Agent Reinforcement Learning (MARL), agents are often trained under concurrent participation, while in many tasks some agents act earlier and leave task-relevant information that becomes useful to agents participating later. We study this setting as staggered participation (SP), which introduces a cross-time, cross-agent learning dependency because an early action may affect the return through the information it provides and the later policy that uses it. Learning under SP therefore requires both identifying what information is useful for future decisions and learning how later agents should use it. We propose Staggered Participation Learning (SPL), a training-time augmentation that addresses these two parts with prospective acquisition supervision for earlier agents and outcome-supervised receiver learning for later agents. We evaluate SPL across multiple policy-based MARL backbones, environments, and staggered-participation patterns. Across 60 MPE/RWARE backbone setting comparisons, SPL achieves higher observed mean task completion in every case, with an average difference of 14.1%. The gains also extend to eight-agent teams and a physics-based UAV-UGV environment in Isaac Lab, providing evidence across algorithmic, temporal, and embodied settings.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07578 [cs.AI]
  (or arXiv:2610.07578v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07578

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jianglin Qiao [view email]
[v1] Tue, 6 Oct 2026 01:13:24 UTC (10,509 KB)

来源:arXiv:cs.AI · arxiv.org