跳到正文
arXiv:cs.LG· Xiaoyang Cao, Jingqi Li, Zhe Fu, Alexandre M. Bayen·· 4 小时前AI 评分34

LiRA:多智能体强化学习中共享约束的责任分配方法

Who Bears the Burden? Learning Responsibility for Shared Constraints in Multi-Agent Reinforcement Learning

AI 导读

针对多个智能体共享成本预算时惩罚难以分摊的问题,研究者提出 Lagrangian Responsibility Allocation(LiRA),通过学习每个智能体对公共 Lagrange 乘子的责任份额,在社会福利目标下优化有限训练周期内的分配。

正文

View PDF HTML (experimental)

Abstract:When multiple agents share a cost budget, a common Lagrange multiplier can enforce the aggregate constraint but does not determine how its penalty should be allocated across agents. Uniform penalties ignore heterogeneity in the rewards agents sacrifice, while agent-specific multipliers may still rely on the same aggregate cost signal. We introduce Lagrangian Responsibility Allocation (LiRA), which learns each agent's share of a common multiplier by optimizing social welfare over a finite training horizon. The multiplier enforces the aggregate budget, while responsibility shares redistribute its influence without modifying the original rewards or constraints. For convex games under standard regularity conditions, varying these shares induces a smooth family of normalized generalized Nash equilibria in which active constraints remain at their budgets while welfare varies. To optimize responsibility before convergence, we derive a welfare gradient that accounts for both learning updates and the induced change in data distribution. Across CityLearn, MABIM, Harvest, and MetaDrive, spanning 3 to 400 agents, LiRA improves average social welfare by up to 29% over uniform and agent-specific multiplier baselines. Grid and driving costs remain within budget, inventory violations decrease, and Harvest makes more effective use of available budget.
Comments: 20 pages, 2 figures, 4 tables. Project page with code: this https URL
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Science and Game Theory (cs.GT); Multiagent Systems (cs.MA)
Cite as: arXiv:2610.07491 [cs.LG]
  (or arXiv:2610.07491v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07491

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xiaoyang Cao [view email]
[v1] Mon, 5 Oct 2026 22:53:33 UTC (374 KB)

来源:arXiv:cs.LG · arxiv.org