跳到正文
arXiv:cs.LG· Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao, Se-Young Yun, Rahul G. Krishnan·· 3 小时前AI 评分40

Pivot-SD:面向掩码扩散语言模型的高效自蒸馏方法

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

AI 导读

Pivot-SD 是一种面向掩码扩散语言模型(dLM)的离线自蒸馏框架,仅监督去噪过程中信息增益最高的关键承诺(pivots),成功轨迹的 pivots 用交叉熵训练,失败轨迹的 pivots 用定向 unlikelihood 训练。

正文

View PDF HTML (experimental)

Abstract:Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an information-gain metric measuring uncertainty reduction over the remaining masked positions. Pivots from successful trajectories are trained with cross-entropy, and pivots from failed trajectories with targeted unlikelihood, leaving the rest of the failed trajectory untouched. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.
Comments: EMNLP 2026 Main (Oral)
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as: arXiv:2610.03665 [cs.LG]
  (or arXiv:2610.03665v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03665

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sunwoo Hong [view email]
[v1] Fri, 2 Oct 2026 17:37:51 UTC (260 KB)

来源:arXiv:cs.LG · arxiv.org