跳到正文
arXiv:cs.LG· Zhengyu Fang, Seoyeon Hong, Jie Yang, Muyang Li, Koyoshi Shindo, Brandon Joseph Lwowski, Jing Li·· 3 小时前AI 评分32

Latent-MOPD:面向 LLM 的潜在多教师同策略蒸馏

Latent-MOPD: Latent Multi-Teacher On-Policy Distillation

AI 导读

Latent-MOPD 是首个面向 LLM 的表征级多教师同策略蒸馏(OPD)方法,无需额外教师训练即可同时利用专家的预测分布与隐藏状态进行监督。在同族设置下,它在数学、代码、逻辑共 9 个基准上全面超越仅 token、仅表征和均匀平均基线,并在多数基准上超过各单项最佳教师。使用跨族教师时,它在所有基准上均优于单通道基线。

正文

View PDF HTML (experimental)

Abstract:On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialists through both their predictions and the hidden states used to compute them, without additional teacher training. To coordinate representation supervision from multiple specialists, we select late-layer targets according to the teacher-student relationship, bridge unequal hidden widths with a shared projection, and group updates by domain. Each teacher's supervision gradually shifts from hidden states to token predictions, with both channels using the same routed specialist. In our main same-family setting, Latent-MOPD outperforms the token-only, representation-only and uniform-averaging baselines on all nine benchmarks across math, code and logic. With the same parameter count as each teacher, the student also surpasses the per-benchmark best teacher on a majority of these benchmarks. With larger, separately developed cross-family teachers, Latent-MOPD outperforms both single-channel baselines on all benchmarks. A same-family all-layer representation-only control remains stable with domain-pure updates but collapses when teacher domains are interleaved within an update. Our results show that a single student can integrate capabilities from several specialists through both their output distributions and internal representations.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02381 [cs.LG]
  (or arXiv:2610.02381v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02381

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhengyu Fang [view email]
[v1] Thu, 1 Oct 2026 19:01:56 UTC (668 KB)

来源:arXiv:cs.LG · arxiv.org