跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Qiwei Di, Xuheng Li, Kaixuan Ji, Chenggong Zhang, Heyang Zhao, Quanquan Gu·· 5 小时前AI 评分35

理解离策略与在策略蒸馏:两种不同训练目标的故事

Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives

AI 导读

研究对比了在策略蒸馏(OPD)与离策略蒸馏在多教师序列蒸馏中的差异:前向 KL 散度产生加权算术混合目标,反向 KL 散度产生归一化加权几何聚合目标,两者分别对应离策略与在策略反馈下的学习算法,并在表格设定下建立了对数遗憾界。分析表明,反向 KL 在无信息反馈下更能保留置信专家的偏好,但对给正确答案分配极低概率的教师更敏感,其 token 级条件分布依赖续写分布,长程下可能偏向错误前缀。

正文

View PDF HTML (experimental)

Abstract:On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers. Forward Kullback--Leibler (KL) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate. We develop algorithms that learn these targets under off-policy and on-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation. By analyzing these aggregation targets, we identify mechanisms that help explain both the benefits and fragility of OPD. Relative to forward KL, reverse KL can better retain a confident expert's preferences under uninformative feedback, but is more sensitive to teachers that assign very low probabilities to correct responses. Its token-level conditionals also reveal a dependence on continuation distributions that can favor incorrect prefixes over long horizons.
Comments: 68 pages, 4 figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as: arXiv:2609.38666 [cs.LG]
  (or arXiv:2609.38666v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38666

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Qiwei Di [view email]
[v1] Tue, 29 Sep 2026 23:47:58 UTC (144 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org