arXiv:cs.CL· Yixuan Tang, Yi Yang·· 3 小时前AI 评分45
On-Policy Distillation 教出新技能而非新知识
On-Policy Distillation Teaches New Skills but Not New Knowledge
AI 导读
一项受控合成实验发现,reverse-KL 的 on-policy distillation(OPD)能跨未见推理结构传递组合技能,却几乎不传递事实知识。拆解蒸馏配方显示,把 reverse KL 换成 forward KL 可恢复事实迁移,而 student rollouts 则专门提升多步推理执行能力。近期事实 QA 与竞赛数学实验呈现同样不对称:推理增益明显,但模型参数化知识未扩张。
正文
Abstract:On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently controls the teacher's additional facts, compositional skill, or both. Across four models from three families, reverse-KL OPD reliably transfers compositional skill across unseen reasoning structures, but transfers minimal factual knowledge. Decoupling the distillation recipe reveals the source of this asymmetry: replacing reverse KL with forward KL restores factual transfer, whereas student rollouts specifically improve the execution of multi-step reasoning. Experiments on recent factual QA and competition mathematics show a similar asymmetry under reverse-KL OPD, yielding notable reasoning gains without factual memory expansion. Together, these results demonstrate that on-policy distillation does not expand a model's parametric knowledge, but instead teaches it to organize and compose the knowledge it already possesses.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.09639 [cs.CL] |
| (or arXiv:2610.09639v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09639 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yixuan Tang [view email]
[v1]
Wed, 7 Oct 2026 08:15:10 UTC (313 KB)
来源:arXiv:cs.CL · arxiv.org