跳到正文
arXiv:cs.LG· Purab Seth, Neil Shah, Ishaan Sinha, Kunal Jha, Samuel J. Gershman, Max Kleiman-Weiner, Wilka Carvalho·· 7 小时前AI 评分42

任务多样性带来系统性迁移却抑制持续强化学习:Banyan 域揭示持续 RL 瓶颈

Task diversity produces systematic transfer but inhibits continual reinforcement learning

AI 导读

研究者提出 GPU 加速的持续强化学习域 Banyan,可分别控制地图布局、交互对象和子目标依赖层级三个任务维度。实验发现,沿每个维度增加多样性会带来系统性迁移,但多样度过高会抑制智能体继续适应新任务分布,成功率陷入平台期,同时旧任务表现仍持续提升。该现象在多种持续学习算法、记忆架构与架构规模以及 Kinetix 物理控制域中均出现,代码已开源。

正文

View PDF HTML (experimental)

Abstract:Continual reinforcement learning (RL) aims to produce agents that never stop adapting to new tasks. A key question is how this interacts with the diversity of tasks an agent experiences. Prior work has shown that training on many diverse tasks leads to agents with strong zero-shot and in-context adaptation. However, this work evaluated agents after they'd stopped learning, i.e. with frozen weights. How task diversity affects an agent's ability to continue learning over a sequence of distribution shifts remains unclear. We introduce Banyan, a GPU-accelerated continual RL domain where one can parametrically control three independent axes that define a task: the map layouts an agent must navigate, the objects it must interact with, and the hierarchical structures of sub-goal dependencies. We find that increasing diversity along each axis induces systematic transfer -- that is, agents begin training on a new task distribution near the performance attained on the previous one, even when the shift changes the structure of the optimal policy. While increasing diversity improves systematic transfer, we find that too much diversity inhibits a learner's ability to continue adapting to new task distributions. As diversity increases, learners plateau in the success rate they achieve on new tasks, yet continue improving on old tasks -- even without further exposure to them. We find this phenomenon manifests across continual learning algorithms, memory architectures, architecture sizes, and in Kinetix -- a physics-based control domain. We release Banyan as a domain for running controlled experiments that study continual RL in the many-tasks regime. Code is available at this https URL.
Comments: 27 pages, 17 figures. v2 adds Kinetix, transformer, and continual-learning-method experiments. Code: this https URL
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2606.00880 [cs.LG]
  (or arXiv:2606.00880v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2606.00880

arXiv-issued DOI via DataCite

Submission history

From: Wilka Carvalho [view email]
[v1] Sat, 30 May 2026 20:31:25 UTC (766 KB)
[v2] Tue, 6 Oct 2026 14:06:33 UTC (2,229 KB)

来源:arXiv:cs.LG · arxiv.org