arXiv:cs.AI· Zipeng Wang, Xinpeng Dong, Yuefan Wang, Pingchen Lu, Xian Wei, Kun Kuang, Fei Wu, Zhongxiang Dai, Min Zhang·· 4 小时前
MetaOPD:面向在线策略蒸馏的元学习 Token 加权框架
MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation
AI 导读
MetaOPD 提出一种双层优化框架,在在线策略蒸馏(OPD)中联合训练学生模型与轻量级 token 加权网络,通过虚拟更新后的验证损失优化权重网络,使预测信号到 token 权重的映射随学生模型共同演化。在六个数学推理和三个域外数据集上,0.6B 学生模型的 Avg@8/Pass@8 较 OPD 提升 1.99/5.97 个百分点,1.7B 学生模型提升 2.25/6.41 个百分点。
正文
Abstract:On-policy distillation (OPD) trains a student on its own generated responses using token-level teacher supervision. However, uniform weighting overlooks differences in token learning value, while existing weighting methods rely on predefined mappings from prediction signals to token weights. These mappings are not learned from the effectiveness of the resulting student updates, limiting their ability to adapt to evolving learning needs. In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network. The inner objective updates the student through weighted OPD, while the outer objective optimizes the weighting network using validation loss on reference solutions after a virtual student update. Differentiating through this update connects weighting decisions to their effects on post-update performance, allowing the mapping from prediction signals to token weights to evolve alongside the student. Experiments on six mathematical reasoning and three out-of-domain datasets, covering two student scales and seven baselines, demonstrate the effectiveness of MetaOPD, with Avg@8/Pass@8 gains over OPD of 1.99/5.97 percentage points for the 0.6B student and 2.25/6.41 points for the 1.7B student.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11989 [cs.AI] |
| (or arXiv:2610.11989v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11989 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Min Zhang [view email]
[v1]
Thu, 8 Oct 2026 13:58:39 UTC (3,448 KB)
来源:arXiv:cs.AI · arxiv.org