跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du, Longbo Huang·· 14 小时前AI 评分37

OMAF:一步在线多智能体流策略,训练提速且样本效率提升 10.5 倍

Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies

AI 导读

研究者提出在线 MARL 框架 OMAF,用基于 Transformer 的流策略实现一步动作生成,替代扩散策略的迭代采样。该方法通过近似路径得分代理与 softmax Q 值联合优化,在 MPE 和 MAMuJoCo 的 10 项任务上回报最高提升 3.4 倍、样本效率提升 10.5 倍。

正文

View PDF HTML (experimental)

Abstract:Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. Generative policies, particularly diffusionbased policies, can faithfully capture complex and multimodal behaviors, but costly iterative sampling hinders their scalability in online multi-agent settings. We propose an Online MARL framework via one-step Flow model (OMAF) that combines expressive generative policies with efficient one-step action generation. OMAF employs a Transformer-based flow policy to capture complex coordination behaviors, while its approximate path score surrogate provides a principled route to synchronized flow policy optimization. To enable stable and sampleefficient learning, we further develop a joint optimization scheme coupling softmax Q-value estimation with a joint flow policy objective for coordinated policy learning. By eliminating iterative sampling, OMAF dramatically reduces training overhead without sacrificing policy expressiveness. Extensive experiments across 10 standard tasks from MPE and MAMuJoCo show that OMAF consistently achieves superior performance, with up to 3.4x higher returns and 10.5x sample efficiency improvement compared with baseline methods. These results validate the effectiveness of OMAF as an expressive and computationally efficient one-step flow policy paradigm for online MARL.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.01882 [cs.LG]
  (or arXiv:2610.01882v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01882

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhuoran Li [view email]
[v1] Thu, 1 Oct 2026 15:36:31 UTC (4,768 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org