arXiv:cs.LG· Shoichiro Takeda, Shin'ya Yamaguchi, Satoshi Suzuki, Yasunori Akagi·· 4 小时前AI 评分39
面向低秩适配的黎曼几何:新度量让 LoRA 更接近全量微调
A Riemannian Geometry for Low-rank Adaptation
AI 导读
研究者提出一种专为 LoRA 定制的黎曼度量,通过等价关系 (B, A) ~ (BG⁻¹, AGᵀ) 诱导的商流形消除冗余方向,并在每次梯度步引入预条件。理论证明该预条件使 LoRA 的权重更新在 LoRA 参数化允许的一阶变化子空间内最接近全量微调梯度方向,且更新后权重矩阵的 Frobenius 范数比常规预条件和无预条件 LoRA 更接近全量微调,语言与视觉领域微调实验验证了其有效性与效率。
正文
Abstract:Low-rank adaptation (LoRA) is widely used as a parameter-efficient fine-tuning technique for pre-trained deep neural networks, which approximates the weight update via full fine-tuning by a low-rank matrix $BA^\top$. This parameterization leads to the equivalence relation $(B, A) \sim (BG^{-1}, AG^\top)$ for any invertible matrix $G$ because $BA^\top = BG^{-1}(AG^\top)^\top$ and thus both pairs yield the same loss value. This relation induces a quotient manifold where matrices $(BG^{-1}, AG^\top)$ for all $G$ are identified, eliminating redundant directions along which the loss value remains unchanged. To respect the geometry of this manifold, the original search space is endowed with a Riemannian metric that is invariant under the equivalence relation. Such a metric induces preconditioning at each gradient step and ensures that each weight update via LoRA changes the loss value, leading to efficient optimization. In this paper, we propose a new Riemannian metric that is specifically tailored to LoRA to close the gap to full fine-tuning at the weight level. We theoretically show that LoRA with our preconditioning induced by this metric satisfies the following two properties at each iteration: (i) The weight update follows the direction closest to the gradient of full fine-tuning within the subspace of first-order weight changes allowed by the LoRA parameterization. (ii) The updated weight matrix is closer in Frobenius norm to that of full fine-tuning than the updated weight matrices of LoRA with conventional preconditioning and without preconditioning. These theoretical insights suggest that our preconditioning makes LoRA better approximate full fine-tuning, thereby leading to more efficient optimization. Experiments show the effectiveness and efficiency of our preconditioning for LoRA on fine-tuning tasks with language and vision domains.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.08049 [cs.LG] |
| (or arXiv:2610.08049v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08049 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shoichiro Takeda [view email]
[v1]
Tue, 6 Oct 2026 09:47:31 UTC (124 KB)
来源:arXiv:cs.LG · arxiv.org