跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Hang Yin, Haozhe Wang, Yuhua Luo, Zhangqi Pan, Xiaoxing Wang, Junchi Yan·· 14 小时前AI 评分33

ChainLoRA:面向 LLM 持续学习的几何保持任务向量合并方法

ChainLoRA: Geometry-Preserving Task Vector Merging for Continual Learning in LLMs

AI 导读

ChainLoRA 是一个无需回放的持续合并框架,基于链式更新的任务向量几何,结合链式更新训练与流后自适应 SVD 合并。训练时初始化和单侧正交代理仅使用最后一个载体,使历史状态占用与正则开销随任务流增长保持恒定;合并时 Adaptive SVD 提取共享载体并通过 Procrustes 适配对齐最新任务。

正文

View PDF HTML (experimental)

Abstract:Continual parameter-efficient fine-tuning for large language models (LLMs) must balance retention of previously acquired knowledge, adaptation to new tasks, and strict parameter budgets. We present \textbf{ChainLoRA}, a replay-free continual merging framework built on chain-updated task-vector geometry. From a parameter-merging perspective, we formulate a geometric view of forgetting through a measurable interaction between task updates, separating directional overlap from coefficient coupling. Building on this view, ChainLoRA combines chain-updated training with post-stream adaptive SVD merging. During training, initialization and a one-sided orthogonality proxy use only the last carrier, keeping their historical-state footprint and regularization overhead constant as the task stream grows. At merging time, Adaptive SVD extracts a shared carrier and aligns it to the latest task through Procrustes adaptation. Our theoretical analysis shows that Procrustes adaptation facilitates geometric approximate separation of shared and task-specific components. The one-sided proxy further bounds inter-task interference. An effective-rank penalty additionally promotes efficient utilization of the task subspace during continual learning. Experiments show that ChainLoRA achieves state-of-the-art performance among the evaluated replay-free methods on the Large and SuperNI benchmarks, while remaining competitive on Standard CL and attaining almost the closest average scores to the evaluated replay-based method across all three benchmarks.
Comments: 18 pages
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG)
ACM classes: I.2.4; F.4.1
Cite as: arXiv:2610.00431 [stat.ML]
  (or arXiv:2610.00431v1 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2610.00431

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Hang Yin [view email]
[v1] Wed, 30 Sep 2026 16:43:34 UTC (355 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org