跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Eleni Nisioti, Andrea Cossu, Kathrin Korte, Sebastian Risi·· 15 小时前AI 评分37

神经进化如何用于持续强化学习:ES 与 GA 的稳定性-可塑性权衡研究

Continual Reinforcement Learning with Neuroevolution

AI 导读

研究对比进化策略(ES)与遗传算法(GA)和当前持续强化学习方法,发现 ES 最稳定地实现稳定性-可塑性权衡,GA 可塑性最强但遗忘更多。ES 找到的解在权重空间中邻域最宽,相邻任务邻域重叠大小与方法权衡相关;通过 novelty search 奖励行为多样性会让 GA 更可塑但更易遗忘。RL 中常见的可塑性丧失症状未出现在神经进化中。

正文

View PDF HTML (experimental)

Abstract:Despite many studies about causes and remedies of plasticity loss in Reinforcement Learning (RL) under continual task changes, no RL method has yet consistently achieved a good balance between adaptation and forgetting. Here we turn to an alternative optimization paradigm, neuroevolution (NE): algorithms that search directly in weight space through mutation and selection over a population of neural networks. Across a wide array of environments and environmental changes, with policies ranging from a few hundred parameters to million-parameter networks, we compare evolution strategies (ES) and genetic algorithms (GAs) against state-of-the-art continual RL variants and population-based RL. ES most consistently achieves a good stability-plasticity trade-off, while the GA is the most plastic method but forgets more than ES. To explain this, we study the return landscape around each method's solutions. ES finds the widest neighborhoods, i.e.\ regions of weight space in which perturbed policies still solve the task, and the size of the overlap between the neighborhoods of consecutive tasks correlates with a method's stability-plasticity trade-off. Rewarding behavioral diversity in a GA through novelty search makes the population even more plastic, at the cost of forgetting. Finally, symptoms of plasticity loss commonly reported in RL do not transfer to NE. Overall, these results establish NE as a competitive alternative to RL under continual task changes, and suggest that training under perturbations in weight space may be a useful mechanism for continual learning more broadly.
Subjects: Neural and Evolutionary Computing (cs.NE); Machine Learning (cs.LG)
Cite as: arXiv:2610.01583 [cs.NE]
  (or arXiv:2610.01583v1 [cs.NE] for this version)
  https://doi.org/10.48550/arXiv.2610.01583

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Andrea Cossu [view email]
[v1] Thu, 1 Oct 2026 12:39:12 UTC (11,952 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org