跳到正文
arXiv:cs.CL· Haoxuan Luo, Jameson Sandler, Ferdinando Fioretto·· 4 小时前AI 评分37

投机解码中的验证器跳过:从逐位置置信度到前缀调度

From Positionwise Confidence to Prefix Scheduling: Verifier Skipping in Speculative Decoding

AI 导读

研究提出投机解码中的"验证器跳过"策略,即直接提交选定的草稿前缀、跳过目标模型验证,以缓解验证瓶颈。

正文

View PDF HTML (experimental)

Abstract:Speculative decoding is a leading technique to reduce the cost of autoregressive generation by using a small drafter to propose several tokens, which are then verified in parallel by a larger target model. Speculative diffusion decoding (SDD) further removes sequential drafting by generating every position in a draft block in parallel with a discrete diffusion model. However, SDD still invokes the target on every block, leaving verification as a potential bottleneck. This paper recognizes that this creates a new control handle: whether to invoke the verifier at all. Thus, we study verifier skipping, a lossy policy that commits a selected draft prefix directly, and ask which confidence signal should schedule it. Interestingly, our study finds that better token predictors need not yield better schedulers: skips require contiguous high-confidence prefixes, while short skips can induce additional drafting rounds. To study this mismatch, we compare raw confidence with learned marginal and conditional survival scores under the same policy, using Strict SDD, lenience, and top-$k$ acceptance as baselines. On HumanEval with DiffuCoder-7B-Instruct and Qwen3-32B, all three confidence signals save $9.6\%$ to $13.5\%$ of verifier calls at the same observed pass@1 as Strict SDD. Surprisingly, raw confidence saves the most; marginal survival has higher positionwise AUROC than raw confidence at most positions, yet neither learned signal dominates online. Our analysis shows that verifier skipping is a useful new lossy axis and, surprisingly, its key challenge is prefix scheduling rather than token prediction alone.
Comments: Accepted at UncertaiNLP 2026 (non-archival). 14 pages, 6 figures
Subjects: Cryptography and Security (cs.CR); Computation and Language (cs.CL)
Cite as: arXiv:2608.14787 [cs.CR]
  (or arXiv:2608.14787v2 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2608.14787

arXiv-issued DOI via DataCite

Submission history

From: Haoxuan Luo [view email]
[v1] Fri, 14 Aug 2026 18:00:08 UTC (151 KB)
[v2] Thu, 1 Oct 2026 23:52:34 UTC (151 KB)

来源:arXiv:cs.CL · arxiv.org