跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Siqi Zhu, Suozhi Huang, Kaixuan Zhang, Yuheng Yang, Zhanyang Jin, Yihang Sun, Jiaxuan You·· 14 小时前AI 评分39

从梯度到能力:理解多教师同策略蒸馏

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

AI 导读

研究以 Qwen3-1.7B 和四个同源 RL 教师为对象,分析多教师同策略蒸馏(MOPD)中教师信号如何影响参数变化。发现损失平均会隐式加权响应,Adam 一阶矩使教师间参数更新余弦相似度达 0.83,BF16 舍入掩盖了约 97% FP32 主权重与初始化的差异。

正文

View PDF HTML (experimental)

Abstract:Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97\% of FP32 master weights differ from initialization, but only 7--11\% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02179 [cs.LG]
  (or arXiv:2610.02179v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02179

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Siqi Zhu [view email]
[v1] Thu, 1 Oct 2026 17:57:44 UTC (1,444 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org