跳到正文
arXiv:cs.CL· Yunxiao Zhao, Changxiao Cai·· 3 小时前AI 评分40

通过直接最小化期望解码轮数训练并行投机草稿模型

Training Parallel Speculative Draft Models by Directly Minimizing Expected Decoding Rounds

AI 导读

研究者提出期望解码轮数(EDR)目标,将投机解码建模为马尔可夫奖励过程,用状态占用率加权局部拒绝代价,且不引入额外超参数。该方法可推导出精确的时序差分梯度,支持从目标模型 rollout 进行无偏随机优化。用 EDR 微调 DSpark 和 DFly 两个 SOTA 草稿模型后,在数学推理、代码生成和对话共九项基准上持续提升平均接受长度,并优于现有训练目标。

正文

View PDF HTML (experimental)

Abstract:Speculative decoding accelerates large language model inference by using a low-cost draft model to propose tokens that the full-size target model verifies in parallel. Parallel and semi-autoregressive (semi- AR) drafters improve drafting efficiency by proposing an entire block in a single forward pass, but training them raises a new difficulty: the draft distribution for a given position depends on where the decoding round starts, and where rounds start depends on how many tokens earlier rounds accepted. Existing training objectives typically rely on block-local surrogates that ignore this cross-round coupling, and therefore do not directly optimize the global decoding efficiency. In this work, we develop a theoretical framework for training and evaluating these drafters by representing speculative decoding as a Markov reward process. This formulation yields the Expected Decoding Rounds (EDR) objective, which weights local rejection costs by state occupancies and exactly equals the expected number of decoding rounds. Unlike prior surrogate objectives, EDR introduces no auxiliary hyperparameters. We then derive an exact temporal-difference gradient that supports unbiased stochastic optimization from target-model rollouts. The same framework also yields an exact offline evaluator for round counts, enabling paired drafter comparisons on shared target rollouts without running speculative decoding. Finetuning two state-of-the- art drafters, DSpark and DFly, with EDR consistently improves mean accepted length and outperforms existing training objectives across nine benchmarks spanning math reasoning, code generation, and chat.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (stat.ML)
Cite as: arXiv:2610.10411 [cs.LG]
  (or arXiv:2610.10411v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.10411

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yunxiao Zhao [view email]
[v1] Wed, 7 Oct 2026 16:56:29 UTC (59 KB)

来源:arXiv:cs.CL · arxiv.org