arXiv:cs.LG· Myeung Suk Oh, Zhiyao Zhang, Alvaro Velasquez, Nathaniel D. Bastian, Jia Liu·· 4 小时前AI 评分30
基础模型辅助多智能体强化学习优化无线随机接入网络
Foundation Model-Aided Multi-Agent Reinforcement Learning for Wireless Random Access Network Optimization
AI 导读
研究提出一种基础模型(FM)辅助的 actor-critic 算法,嵌入基于共识的去中心化 MARL 架构,用于提升无线随机接入网络优化的 MARL 效率。作者给出该算法在局部奖励交换与非线性值函数近似下的收敛性分析,证明其收敛阶与传统 MARL(带 critic 模型交换与线性近似)相同。数值结果表明,该 FM 方法显著提升了随机接入网络优化的 MARL 速度。
正文
Abstract:Random access (RA) is one of the most foundational medium access control (MAC) layer scheduling schemes for handling unpredictable data traffic from multiple terminals. While multi-agent reinforcement learning (MARL) has been explored to optimize RA-based wireless networks, its reliance on experience-driven, distributed policy learning incurs significant training overhead for each optimization task, limiting its feasibility in real-world applications. In this work, we propose to leverage a foundation model (FM) to improve MARL efficiency across diverse RA network optimization tasks. Specifically, we design an FM-aided actor-critic algorithm within a consensus-based decentralized MARL architecture and provide its convergence analysis under local reward exchanges and nonlinear value function approximations to show that our algorithm achieves the same convergence order as the conventional MARL with critic model exchanges and linear approximations. Our numerical results show that our FM-based approach significantly enhances MARL speed for RA network optimization.
| Comments: | This paper has been accepted in ACM International Symposium on Theory, Algorithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing (MobiHoc) 2026 |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07550 [cs.LG] |
| (or arXiv:2610.07550v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07550 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Myeung Suk Oh [view email]
[v1]
Tue, 6 Oct 2026 00:27:24 UTC (1,185 KB)
来源:arXiv:cs.LG · arxiv.org