arXiv:cs.LG· Srijith Nair, Atilla Eryilmaz, Jia Liu·· 2 天前AI 评分38
PMF-CL:面向冲突任务的帕累托最小遗忘持续学习算法
PMF-CL: Pareto-Minimal-Forgetting Continual Learner for Conflicting Tasks
AI 导读
研究者提出 PMF-CL 持续学习算法,从多任务学习视角将持续学习重新定义为事后平衡冲突任务,通过寻找帕累托最优解来最小化对旧任务的遗忘。该算法仅需 O(d²) 的静态内存占用(d 为模型参数量),且与任务数 T 无关,在二次损失下可获得精确的事后帕累托最优遗忘保证,在二次上界损失下遗忘可被证明有界。数值实验验证了其收敛到帕累托最优解的理论结论,并持续优于现有方法且占用更少内存。
正文
Abstract:In the literature, many continual learning (CL) algorithms have been proposed to address the issue of catastrophic forgetting in ML models (i.e., learning new tasks leads to the loss of performance on previously learned tasks). Although all CL approaches use some form of memory to retain information about past tasks, a grounded understanding of what information needs to be stored to minimize catastrophic forgetting remains elusive. Recently, it has been recognized that under the strong assumption of the existence of a common global minimizer over all tasks, catastrophic forgetting can be completely avoided. However, in practice, tasks rarely have a common global minimizer, and a certain amount of forgetting is inevitable. In this paper, we propose a foundational reframing of CL as balancing conflicting tasks in hindsight from a multi-task learning (MTL) perspective. The approach is based on finding Pareto-optimal solutions, i.e., the solutions which, by definition, minimally forget the previous tasks in the Pareto sense. We derive a Pareto-minimal-forgetting CL (PMF-CL) algorithm for linear and basis-function regression. Our algorithm naturally extends to loss functions with quadratic upper bounds around their minimizers, despite consuming only a static memory footprint of $O(d^2)$ for $d$ model parameters, independent of the number of sequentially occurring tasks $T$. In the quadratic-loss setting, we obtain exact hindsight Pareto-optimal guarantees on forgetting. In the quadratic-upper-bound setting, forgetting is provably bounded, with the tightness of the forgetting bound being directly determined by the tightness of the quadratic bound on the loss function. Our numerical results validate our theoretical claims of convergence to the Pareto-optimal solution, and demonstrate that we consistently out-perform prior empirical methods while occupying lesser memory.
| Comments: | 30 pages, 6 figures, 4 algorithms |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.19145 [cs.LG] |
| (or arXiv:2605.19145v3 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.19145 arXiv-issued DOI via DataCite |
Submission history
From: Srijith Nair [view email]
[v1]
Mon, 18 May 2026 21:53:49 UTC (1,091 KB)
[v2]
Fri, 29 May 2026 03:14:31 UTC (1,091 KB)
[v3]
Wed, 30 Sep 2026 18:13:16 UTC (2,462 KB)
来源:arXiv:cs.LG · arxiv.org