arXiv:cs.LG(机器学习,全量分类)· Yifan Hu, Luhang Hong, Mingkang Long, Danning Wang, Chengfeng Jia, Rong Su, Junjie Fu, Guanghui Wen·· 9 小时前AI 评分36
MASkillBlender:通过技能混合实现多足式机器人全身协调的分散式多智能体强化学习框架
MASkillBlender: Decentralized Whole-Body Coordination for Multi-Humanoid Loco-Manipulation via Skill Blending
AI 导读
MASkillBlender 是一个分散式多智能体强化学习框架,通过在预训练的单足式机器人技能上学习共享的高层策略,实现多足式机器人的全身协调操作,仅需任务级奖励,无需任务特定的运动参考。该框架还引入基于排列的数据增强策略,并证明在齐次马尔可夫博弈下排列样本保持原样本的策略梯度方向。在两种足式机器人本体上的多个协调任务仿真中,该框架均取得强任务表现。
正文
Abstract:Coordinated multi-humanoid loco-manipulation is promising yet challenging due to high-dimensional whole-body control, decentralized decision making, and scalability. While recent reinforcement learning methods have improved single-humanoid whole-body control, extending them to the multi-humanoid setting remains nontrivial and often requires substantial reward engineering or task-specific design. We propose MASkillBlender, a general multi-agent reinforcement learning framework to achieve decentralized multi-humanoid whole-body coordination. By learning a shared decentralized high-level policy over reusable pre-trained single-humanoid skills, MASkillBlender enables coordinated behaviors using only task-level rewards, without requiring task-specific motion references. To improve learning efficiency, we further introduce a permutation-based data augmentation strategy for homogeneous multi-humanoid systems, and theoretically show that the permuted samples preserve the policy-gradient direction of the original samples under the homogeneous Markov game formulation. We evaluate MASkillBlender on multiple multi-humanoid coordination tasks across two humanoid embodiments. Simulation results demonstrate that the proposed framework consistently achieves strong task performance and enables coordinated behaviors across different tasks and humanoid embodiments.
| Subjects: | Robotics (cs.RO); Machine Learning (cs.LG); Multiagent Systems (cs.MA) |
| Cite as: | arXiv:2610.01102 [cs.RO] |
| (or arXiv:2610.01102v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01102 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yifan Hu [view email]
[v1]
Thu, 1 Oct 2026 05:42:42 UTC (10,744 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org