arXiv:cs.LG· Caleb Chang, Davin Win Kyi, Natasha Jaques, Karen Leung·· 5 小时前AI 评分34
GRID:从异构智能体中学习通用行为
Do as the Romans Do: Learning Universal Behaviors from Heterogeneous Agents
AI 导读
研究者提出社会学习方法 GRID(General Reward Inference and Disentanglement),通过信息瓶颈将每个智能体的奖励函数分解为跨智能体共享的通用奖励和体现个体偏好的特定奖励。
正文
Abstract:Humans often acquire new skills by observing others, since observed behaviors implicitly reveal how to act reasonably in an environment. However, observations drawn from a heterogeneous population introduce conflicting behavioral signals, making it difficult to determine which behaviors are worth imitating. We address this challenge with General Reward Inference and Disentanglement (GRID), a social learning method that extracts universally useful behaviors from a heterogeneous population of demonstrators pursuing different goals. GRID decomposes per-agent reward functions into a general reward, capturing behaviors shared across all agents, and specific rewards, capturing individual preferences and objectives, through an information bottleneck. Training exclusively on the general reward provides a new paradigm of generalist pretraining. It yields a generalist agent that internalizes universal environmental competencies, such as safety and basic task proficiency, without the mode-averaging bias that afflicts standard learning from demonstration techniques. This generalist serves as a strong prior for fine-tuning to downstream tasks, including preferences unseen during training. Experiments across a synthetic basis function decomposition, multi-agent Craftax, continuous control tasks (MuJoCo Gym) and an autonomous driving simulator (Highway-Env) confirm that GRID successfully disentangles reward structure in a semantically meaningful way, outperforms standard learning from demonstration baselines, and enables more efficient and stable specialization.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2606.18537 [cs.LG] |
| (or arXiv:2606.18537v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2606.18537 arXiv-issued DOI via DataCite |
Submission history
From: Caleb Chang [view email]
[v1]
Tue, 16 Jun 2026 23:11:14 UTC (4,086 KB)
[v2]
Fri, 2 Oct 2026 06:29:26 UTC (2,146 KB)
来源:arXiv:cs.LG · arxiv.org