跳到正文
arXiv:cs.LG· Doseok Jang, Jon Ander Campos, Youran Qi·· 3 小时前AI 评分33

LMOPD:词典序多目标同策略蒸馏

Lexicographic Multi-Objective On-Policy Distillation

AI 导读

研究者提出词典序多目标同策略蒸馏(LMOPD),一种按显式优先级整合奖励专精策略的多教师方法:对每个学生 rollout,先选出首个被门控判定存在不足的目标所对应的专家,再对其中心化对数策略修正做局部投影,剔除与更高优先级专家相冲突的分量。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward priority order. This is problematic when trade-offs are asymmetric: conciseness, for example, should not improve at the cost of correctness. We introduce Lexicographic Multi-Objective On-Policy Distillation (LMOPD), a multi-teacher method for integrating reward-specialized policies under explicit priorities. For each student rollout, LMOPD selects the specialist for the first objective whose gate detects a deficiency, then locally projects its centered log-policy correction to remove components that oppose higher-priority specialists. We evaluate 30B-A3B mixture-of-experts transformer models in two- and four-expert settings on three math benchmarks, measuring retained specialist gains. With two experts, LMOPD's point estimates fully retain the accuracy and reasoning-quality gains while acquiring $46.9\%$ of the conciseness gain. With four experts, it retains $\approx90\%$ of both the accuracy gain and reasoning-correctness gain, compared to only $\approx57\%$ by the next best evaluated baseline. Matched four-expertablations show that lexicographic routing outperforms random routing and that projection further strengthens both top-priority capabilities. Across both scales, LMOPD preserves the highest-priority capabilities more effectively than the existing baselines we evaluate, demonstrating the value of explicit priorities for specialist integration.
Comments: 24 pages, 3 figures, 5 tables; includes appendices
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.02359 [cs.LG]
  (or arXiv:2610.02359v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02359

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Doseok Jang [view email]
[v1] Thu, 1 Oct 2026 18:35:35 UTC (3,609 KB)

来源:arXiv:cs.LG · arxiv.org