跳到正文
arXiv:cs.LG· Fernando Martinez, Tao Li, Yingdong Lu, Juntao Chen·· 4 小时前AI 评分28

CASTLE:面向独立多智能体强化学习的反事实语义-社会世界模型

Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models

AI 导读

研究提出 CASTLE,一个离线训练、在线上下文引导的框架,为独立学习的智能体配备局部动力学世界模型与语义-社会世界模型,后者通过反事实模拟器 rollout 学习候选动作的队友与对手响应,两个世界模型在线阶段冻结,仅用局部信息为独立 PPO 策略提供预测 logits 引导。

正文

View PDF HTML (experimental)

Abstract:Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication. Such a stringent information structure renders the conventional reward signal ambiguous. A poor return may result from an ineffective ego action, an incompatible teammate response, or an effective opponent response, yet scalar rewards alone do not reveal which explanation is responsible. We argue that agents can learn more effectively by prospectively comparing the consequences of candidate actions rather than diagnosing failures only from realized returns. We introduce CASTLE (Counterfactual Action-conditioned Semantic Tokens for Local Execution in Decentralized MARL), an offline-training, online-in-context guidance framework with two complementary world models. A Local Dynamics World Model, offline pre-trained over agents' local trajectories, summarizes the agent's local trajectory dynamics and partial observability, while a Semantic-Social World Model predicts compact short-horizon task and social consequences for each candidate ego action. The latter is trained from counterfactual simulator rollouts that expose plausible teammate and opponent responses to alternative actions taken from the same logged rollout state. During online learning and execution, both world models remain frozen and are queried by agents using only locally available information. Their prediction logits provide in-context guidance to an independent PPO policy. Across 30 matched seeds on Tag, Spread, and Adversary in the benchmark multi-particle environments, our proposed CASTLE achieves the highest mean final score among the evaluated methods, exceeding the strongest baseline on each task by 10.67, 6.46, and 0.33 normalized points, respectively.
Subjects: Multiagent Systems (cs.MA); Machine Learning (cs.LG)
Cite as: arXiv:2610.07704 [cs.MA]
  (or arXiv:2610.07704v1 [cs.MA] for this version)
  https://doi.org/10.48550/arXiv.2610.07704

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Fernando Martinez [view email]
[v1] Tue, 6 Oct 2026 03:53:47 UTC (1,045 KB)

来源:arXiv:cs.LG · arxiv.org