LMSYS:Blog(Chatbot Arena 团队)·· 13 小时前AI 评分39
Miles 新增 OPD 支持:Qwen3.5-35B-A3B 自蒸馏实验验证
Blog OPD Support in Miles We recently implemented On-Policy Distillation (OPD) as an important feature in Miles. OPD is now integrated into Miles rollout and training flow, so users can train a student model either solely with... Kaixi Hou & Miles Team July 18, 2026
AI 导读
Miles 将 On-Policy Distillation(OPD)集成进 rollout 与训练流程,支持纯蒸馏和结合 GRPO/PPO 的 RL 增强蒸馏两种模式,并新增稀疏的逐位置候选 token 打分流程,避免 O(R²K) 的密集 payload。
来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org