跳到正文
原文
arXiv:cs.AI(全量分类)· Huiwen Yan, Kyriakos G. Vamvoudakis, Mushuang Liu·· 5 小时前AI 评分27

Meta-MARL:面向交互策略快速适应的元多智能体强化学习框架及自动驾驶应用

Meta-Multi-Agent Reinforcement Learning for Fast Adaptation of Interactive Policies with Applications to Autonomous Driving

AI 导读

研究者提出 meta-MARL 框架,将多智能体强化学习建模为 Markov 博弈,实现跨博弈分布的交互策略快速适应,并定义新解概念 meta-NE,证明其与基于 gradient-play 的 meta-MARL 算法稳定点的等价充分条件。在自动驾驶任务上,该方法比预训练 MARL 基线适应更快。

正文

View PDF HTML (experimental)

Abstract:This paper develops a meta-multi-agent reinforcement learning (meta-MARL) framework to enable fast adaptation of interactive policies in a multi-agent system (MAS). Meta-reinforcement learning (meta-RL) enables agents to rapidly adapt to new tasks/environments using a bi-level optimization mechanism. However, existing meta-RL generally focuses on single-agent systems. Extending these frameworks and algorithms to multi-agent systems poses additional challenges, as tasks are characterized by not only the environment but also agents' strategic interactions. To address these challenges, we model multi-agent reinforcement learning (MARL) problems as Markov games (MGs) and develop a meta-MARL framework for rapid interactive policy adaptation across a distribution of MGs. A new concept, called meta-NE, is defined to describe the desired solution concept in a meta-MARL problem. Sufficient conditions for the equivalence between a meta-NE and a stationary point of the gradient-play-based meta-MARL algorithm are established. Our evaluation on autonomous-driving tasks demonstrates that the proposed meta-MARL method achieves faster adaptation than pretrained MARL baselines, validating the effectiveness of our framework.
Subjects: Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Systems and Control (eess.SY)
Cite as: arXiv:2610.00705 [cs.AI]
  (or arXiv:2610.00705v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.00705

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Huiwen Yan [view email]
[v1] Wed, 30 Sep 2026 20:52:19 UTC (2,222 KB)

来源:arXiv:cs.AI(全量分类) · arxiv.org