跳到正文
arXiv:cs.LG· Yu-Du Feng, Niels M\"undler-Sasahara, Mark Vero, Martin Vechev·· 2 天前AI 评分39

利用指令微调与模型合并实现推理模型适配

Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation

AI 导读

研究提出用标准指令微调加模型合并的方式,将推理语言模型适配到新领域:先做指令微调,再按合适比例与原始推理模型合并,以恢复目标域的推理行为。在 4 个推理模型上针对代码与文本摘要任务评测,目标任务性能最高提升 11.0%,分布外分数平均仅下降 0.7%,单模型适配成本低于 10 美元。

正文

View PDF HTML (experimental)

Abstract:Reasoning language models (RLMs) demonstrate impressive performance by leveraging test-time compute in the form of reasoning tokens. However, this behavior makes adapting RLMs to new domains challenging and expensive. The reason is that further training can disturb the learned behavior and degrade model performance. This makes it difficult to leverage supervised fine-tuning data with human-written solutions: although it contains high-quality annotations, it lacks reasoning tokens. In this work, we show how, despite this challenge, such data can be used efficiently for RLM adaptation. For this, we first use standard instruction tuning. Next, we leverage model merging to combine the instruction-tuned model with the original RLM, picking the merging ratio such that the resulting model's reasoning behavior on the target domain is recovered. We evaluate our method across four RLMs on coding and text summarization tasks, where it improves target-task performance by up to $11.0\%$ while preserving reasoning behavior and limiting the out-of-distribution score degradation to on average $0.7\%$. Importantly, our adaptations are efficient and economical, costing less than USD $\$10$ per model.
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as: arXiv:2607.14895 [cs.LG]
  (or arXiv:2607.14895v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2607.14895

arXiv-issued DOI via DataCite

Submission history

From: Niels Mündler-Sasahara [view email]
[v1] Thu, 16 Jul 2026 12:11:25 UTC (1,126 KB)
[v2] Thu, 1 Oct 2026 16:18:04 UTC (252 KB)

来源:arXiv:cs.LG · arxiv.org