arXiv:cs.LG· Xiaofei Yin, Tong Chu, Jiyuan Fu, Jun Lan, Shuheng Zhou, Huijia Zhu·· 3 小时前AI 评分32
SF-MOPD:慢-快多教师在线策略蒸馏实现能力保持
Slow-Fast Multi-Teacher On-Policy Distillation for Capability Preservation
AI 导读
研究者提出慢-快多教师在线策略蒸馏(SF-MOPD),通过将学生模型与其实时更新的指数移动平均"慢模型"耦合,在 log-probability 空间剔除使快模型偏离慢模型的分量,缓解多教师在线策略蒸馏(MOPD)中的能力干扰。实验表明,SF-MOPD 在多个模型规模上有效提升专业多模态能力,并降低通用能力基准上的平均退化,持续优于 vanilla MOPD。
正文
Abstract:Foundation multimodal large language models are designed to support a broad spectrum of capabilities across diverse domains. Multi-teacher on-policy distillation (MOPD) provides an effective framework for consolidating domain-specific expertise into a single student model. However, MOPD training gradually drives the student away from its initialization model, and general capabilities decline as the displacement grows, resulting in capability interference. A direct remedy is constraining the student toward its initialization, but this suppresses the acquisition of domain expertise as well. We propose Slow-Fast Multi-Teacher On-Policy Distillation (SF-MOPD), which couples a fast model, the current student updated directly by each teacher, with a slow model, an exponential moving average of the student. The slow model absorbs the learning signal gradually, serving as a moving capability reference that fuses the general foundation with confirmed domain expertise. For each teacher, SF-MOPD computes the teacher-induced update in log-probability space and removes only the component that pushes the fast model further away from the slow model, while retaining aligned and orthogonal components. Experiments across multiple model scales demonstrate that SF-MOPD effectively mitigates capability interference, enhances specialized multimodal capabilities, and reduces the average degradation on general-capability benchmarks, consistently outperforming vanilla MOPD.
| Comments: | 5 pages, 2 figures |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02324 [cs.LG] |
| (or arXiv:2610.02324v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02324 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xiaofei Yin [view email]
[v1]
Thu, 1 Oct 2026 18:00:27 UTC (1,271 KB)
来源:arXiv:cs.LG · arxiv.org