跳到正文
arXiv:cs.AI· Zhiqiang Pang, Zihong Sun, Qi Xie, Jun Shu, Deyu Meng, Zongben Xu·· 3 小时前

在哪里适配很重要:面向能力保留的层选择性微调

Where to Adapt Matters: Layer-Selective Fine-Tuning for Capability Retention

AI 导读

研究者提出 LS-LoRA,只把可训练的 LoRA 适配器放在输入—输出余弦相似度较低、经验 Fisher 信息较高的 Transformer 层,而非全部层。在数学推理与代码生成任务上,LS-LoRA 在提升目标任务平均表现的同时,比标准全层 LoRA 保留更多常识推理能力。该方法用前向即可计算的余弦相似度替代需反向计算的 Fisher 分数来排序层敏感度。

正文

View PDF HTML (experimental)

Abstract:Parameter-efficient fine-tuning (PEFT) enables large language models (LLMs) to adapt to specialized tasks, but often at the cost of degrading general capabilities acquired during pretraining. Existing approaches primarily mitigate this trade-off through data replay or regularization, relying on additional data or explicit optimization constraints. We instead focus on a different question: where should adaptation be applied? We find that fine-tuning different Transformer layers produces different target-task gains and degrees of capability degradation, suggesting that not all layers are equally suitable for adaptation. To characterize this difference, we use layer-wise empirical Fisher information to measure target-task sensitivity. However, computing Fisher scores requires backward computation and becomes increasingly expensive for large models. We therefore introduce input--output cosine similarity as a lightweight, forward-only proxy for ranking layer sensitivity. Across models and tasks, layers with lower input--output similarity consistently exhibit higher empirical Fisher scores. Building on this observation, we propose Layer-Selective LoRA (LS-LoRA), which places trainable LoRA adapters only in layers with low input--output similarity. Experiments on mathematical reasoning and code generation show that LS-LoRA improves average target-task performance while retaining substantially more commonsense reasoning capability than standard all-layer LoRA, demonstrating that carefully choosing where to adapt can provide a simple and effective way to balance target-task adaptation and general capability retention.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11620 [cs.AI]
  (or arXiv:2610.11620v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.11620

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhiqiang Pang [view email]
[v1] Thu, 8 Oct 2026 09:59:16 UTC (363 KB)

来源:arXiv:cs.AI · arxiv.org