arXiv:cs.LG(机器学习,全量分类)· Hang Yin, Haozhe Wang, Yuhua Luo, Zhangqi Pan, Xiaoxing Wang, Junchi Yan·· 14 小时前AI 评分33
ChainLoRA:面向 LLM 持续学习的几何保持任务向量合并方法
ChainLoRA: Geometry-Preserving Task Vector Merging for Continual Learning in LLMs
AI 导读
ChainLoRA 是一个无需回放的持续合并框架,基于链式更新的任务向量几何,结合链式更新训练与流后自适应 SVD 合并。训练时初始化和单侧正交代理仅使用最后一个载体,使历史状态占用与正则开销随任务流增长保持恒定;合并时 Adaptive SVD 提取共享载体并通过 Procrustes 适配对齐最新任务。
正文
Abstract:Continual parameter-efficient fine-tuning for large language models (LLMs) must balance retention of previously acquired knowledge, adaptation to new tasks, and strict parameter budgets. We present \textbf{ChainLoRA}, a replay-free continual merging framework built on chain-updated task-vector geometry. From a parameter-merging perspective, we formulate a geometric view of forgetting through a measurable interaction between task updates, separating directional overlap from coefficient coupling. Building on this view, ChainLoRA combines chain-updated training with post-stream adaptive SVD merging. During training, initialization and a one-sided orthogonality proxy use only the last carrier, keeping their historical-state footprint and regularization overhead constant as the task stream grows. At merging time, Adaptive SVD extracts a shared carrier and aligns it to the latest task through Procrustes adaptation. Our theoretical analysis shows that Procrustes adaptation facilitates geometric approximate separation of shared and task-specific components. The one-sided proxy further bounds inter-task interference. An effective-rank penalty additionally promotes efficient utilization of the task subspace during continual learning. Experiments show that ChainLoRA achieves state-of-the-art performance among the evaluated replay-free methods on the Large and SuperNI benchmarks, while remaining competitive on Standard CL and attaining almost the closest average scores to the evaluated replay-based method across all three benchmarks.
| Comments: | 18 pages |
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG) |
| ACM classes: | I.2.4; F.4.1 |
| Cite as: | arXiv:2610.00431 [stat.ML] |
| (or arXiv:2610.00431v1 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00431 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hang Yin [view email]
[v1]
Wed, 30 Sep 2026 16:43:34 UTC (355 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org