跳到正文
arXiv:cs.LG· Jingyuan Huang, Zuming Huang, Yucheng Shi, Zhongzhi Li, Xiaoming Zhai, Wei Chu, Ninghao Liu·· 6 小时前AI 评分40

Train4Merge:RL 与 SFT 教师用于 OPD 模型合并的受控单教师对比研究

Train4Merge: A Controlled Single-Teacher Study of RL vs. SFT Teachers for OPD-Based Model Merging

AI 导读

一项受控单教师研究对比了 SFT 与 RL 训练出的教师模型在基于 OPD 的模型合并中的表现。教师与学生均以 Qwen3.5-9B 初始化,在 Agentic、Reasoning、Perception 三个领域,RL 教师引导学生均优于 SFT 教师,最佳 checkpoint 分别高出 4.27、1.50、0.86 个百分点。

正文

View PDF HTML (experimental)

Abstract:Domain experts trained from a shared checkpoint can transfer their specialized capabilities to a single student through on-policy distillation (OPD). Existing research primarily focuses on improving this merging process, while the algorithms used to train the experts have received limited systematic comparison. We investigate which training algorithm produces teachers better suited to OPD through controlled single-teacher comparisons of supervised fine-tuning (SFT) and reinforcement learning (RL) across Agentic, Reasoning, and Perception. Teachers and students share the same Qwen3.5-9B initialization, and the two teacher types are compared at similar task performance. Our experiments show that RL teachers yield stronger students and higher recovery of teacher performance gains across all three domains. At their best checkpoints, RL-guided students outperform SFT-guided students by 4.27, 1.50, and 0.86 percentage points, respectively. In Agentic, the best SFT-guided student recovers only 44.44% of its teacher's performance gain over the base model, whereas the best RL-guided student recovers 115.00%, surpassing its teacher. Our further analysis shows that RL teachers undergo smaller parameter displacements from the shared initialization than SFT teachers. These findings support the hypothesis that RL teachers' smaller departures from the student's starting point facilitate learning through OPD, resulting in stronger students.
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2609.32303 [cs.AI]
  (or arXiv:2609.32303v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2609.32303

arXiv-issued DOI via DataCite

Submission history

From: Jingyuan Huang [view email]
[v1] Sat, 26 Sep 2026 06:59:03 UTC (130 KB)
[v2] Wed, 7 Oct 2026 02:15:02 UTC (134 KB)

来源:arXiv:cs.LG · arxiv.org