arXiv:cs.LG· Nicolas Zucchet, Scott W. Linderman·· 3 小时前AI 评分44
散度如何控制知识蒸馏中的熵
Divergence controls entropy in distillation
AI 导读
研究从熵的角度分析知识蒸馏中学生模型熵与散度的关系,证明前向 KL 会将学生模型熵推高至教师模型之上,交叉熵训练作为特例可定量验证。反向 KL 则压缩熵直至师生差距过大,on-policy 蒸馏的低熵来自 token 级反向 KL 而非采样方式。散度因此充当隐式熵正则化器,在自蒸馏中作用最明显。
正文
Abstract:Distillation has become a core primitive of large language model training, but its properties are not yet well understood. We take an entropic perspective, studying how the entropy of the student depends on the data and the divergence that define the distillation objective. We prove that forward KL inflates the entropy of the student above that of the teacher. Since cross-entropy training is a special case, this yields an identity that we verify quantitatively in pretraining and supervised finetuning. Other divergences come with no such guarantee: reverse KL deflates entropy until the gap between student and teacher gets too large, and interpolating between the two changes entropy smoothly early in training but abruptly at convergence. The lower entropy of on-policy distillation comes from token-level reverse KL, not from on-policy sampling. The divergence therefore acts as an implicit entropy regularizer, whose role is clearest in self-distillation: as conditioning on privileged information deflates entropy, the divergence hyperparameters that work best are those that compensate for it.
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.03529 [cs.LG] |
| (or arXiv:2610.03529v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03529 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Nicolas Zucchet [view email]
[v1]
Fri, 2 Oct 2026 16:15:53 UTC (2,283 KB)
来源:arXiv:cs.LG · arxiv.org