跳到正文
arXiv:cs.LG· Wentao Lu, Jesse Clark, Tianyu Zhu·· 3 小时前AI 评分41

冻结塔式转换保留生成能力:MoE LLM 的低预算 AR 到扩散模型转换

Context-Tower Conversion Preserves Generation While Freezing Retains Knowledge: Low-Budget AR-to-Diffusion Conversion of MoE LLMs

AI 导读

研究对比了同一 30B MoE 父模型的两种 AR 到扩散模型转换方案:原地更新部分权重与冻结因果副本经交叉注意力条件化。在 1B 训练 token 预算下,冻结塔模型 HumanEval pass@10 达 71.60,原地模型仅 6.19,提升 11.6 倍,同时保留父模型 95% 的 GSM8K 和 99% 的 MMLU-Pro 分数。

正文

View PDF HTML (experimental)

Abstract:Converting a pretrained autoregressive (AR) model to a diffusion language model (dLLM) enables parallel generation without pretraining a new model. Published conversion methods differ by roughly three orders of magnitude in training data and have not been compared under a common protocol. We compare two conversions of the same 30B Mixture-of-Experts (MoE) parent, holding the corpus, supervised-token budget, trainable parameter set and evaluation harness fixed, each under its own training recipe. The in-place model updates a subset of the parent's weights using denoising and representation-alignment losses; the frozen-tower model instead conditions through cross-attention on a frozen causal copy of the parent. With 1B training tokens, the frozen-tower model scores 71.60 on HumanEval pass@10 against 6.19 for the in-place model, an 11.6x improvement. At the same budget it also keeps 95% of the parent's GSM8K score and 99% of its MMLU-Pro score. A dense-parent experiment reproduces the HumanEval separation. Within the two-tower design at about 500M tokens, freezing the context tower retains substantially more MMLU-Pro performance than training it, while both give similar observed HumanEval scores. Our theoretical analysis establishes that both conversion classes contain an exact sampler for the AR parent under a hard attention mask and left-to-right commitment of one position per round. Under a shared loss, freezing removes the gradient contribution through the context states. Furthermore, evaluation protocol substantially affects a published 500B-token conversion's scores in both directions across tasks, while its AR parent's scores vary by less than three points, so comparing dLLMs needs a common protocol. These results show that, in the tested low-budget regime, the frozen-tower configuration retains substantially more of the parent's generation performance than in-place conversion.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02657 [cs.LG]
  (or arXiv:2610.02657v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02657

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Wentao Lu [view email]
[v1] Fri, 2 Oct 2026 01:26:57 UTC (77 KB)

来源:arXiv:cs.LG · arxiv.org