arXiv:cs.LG· Saptarshi Nath, Inish M. D'Souza, Antonio Carta, Soheil Kolouri, Andrea Soltoggio·· 3 小时前AI 评分36
如何在终身强化学习中寻找并复用策略以实现持续适应
How to Find and Reuse Policies for Continuous Adaptation in Lifelong Reinforcement Learning
AI 导读
研究提出 Adaptive Mask Selection and Composition(AMSC),通过非参数 Wasserstein 任务嵌入从状态-动作-奖励样本中在线估计任务相似度,并用 z-score 归一化的 sparsemax 生成可变支撑集,周期性选择并加权已有策略构成新任务的先验。
正文
Abstract:In lifelong reinforcement learning, retaining previously learned policies is not sufficient for effective transfer to a new task. Useful knowledge may be distributed across several prior policies, and its relevance may change as the learner acquires experience. One hypothesis is that task similarity can be effectively used in a continual learning setting to find and combine previously learned policies. To test it, Adaptive Mask Selection and Composition (AMSC) is designed to estimate similarity from online experience via non-parametric Wasserstein task embeddings from state-action-reward samples. The z-score-normalized sparsemax of the similarity scores are used to derive a variable-size support to periodically choose and weight policies to form a prior when learning a new task. On CT-graph and MiniGrid, AMSC achieves higher mean performance and forward transfer than the evaluated modular composition baselines while exhibiting no forgetting. Results on Continual World suggest that identifying relevant prior knowledge and determining its layer-specific composition may require additional layer-specific tuning. Ablations show that selecting relevant sources and determining how strongly to reuse them are central to these gains. Independently measured pairwise transfer is also positively associated with task-embedding similarity. These results indicate that task similarity can be an effective criterion to select and weight specific knowledge for reuse in lifelong reinforcement learning.
| Comments: | Code is available at this https URL |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO) |
| Cite as: | arXiv:2610.03119 [cs.LG] |
| (or arXiv:2610.03119v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03119 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Saptarshi Nath [view email]
[v1]
Fri, 2 Oct 2026 10:37:52 UTC (564 KB)
来源:arXiv:cs.LG · arxiv.org