跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Jungseob Lee, Sugyeong Eo·· 14 小时前AI 评分49

CAST:用单次块草稿器构建成本感知的推测树,加速 LLM 推理最高 43%

CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters

AI 导读

CAST(Cost-Aware Speculative Trees)将块草稿器一次前向打分出的多个候选打包成树,用目标模型单次验证,最高比标准链式推测解码快 43%。它按"下一个候选的预期收益是否超过验证耗时"动态决定树宽,无需扫描调参,在三个 GPU 世代、两个模型族、五个领域的八种设置中均更快。CAST 证明在贪心和采样解码下均不改变目标输出分布,代码已开源。

正文

View PDF HTML (experimental)

Abstract:Speculative decoding accelerates large language model inference by drafting future tokens cheaply and verifying them with the target model in parallel. Block drafters score a whole block of future tokens in one forward pass, yet standard decoding verifies only the top-scoring chain and discards the other candidates. Because these candidates are already scored, verifying more of them adds target computation but no extra drafting. We introduce CAST (Cost-Aware Speculative Trees), which packs these candidates into a tree and verifies it in a single target pass, leaving the target model, drafter weights, and decoding rule untouched. To decide how wide the tree should be, CAST adds candidates while the expected gain from the next one outweighs the verification time it adds. The width therefore adapts to each deployment from a latency measurement, without sweeping over widths. We evaluate CAST across five domains on three GPU generations and two model families. At its predicted width, CAST is faster than the standard chain in all eight settings, by up to 43%. We also find that the best width depends strongly on the deployment. Where verification cost jumps at a kernel boundary, a 128-token tree is only 2% faster than the standard chain, whereas the tree at the predicted width is 20% faster. Furthermore, we prove that CAST leaves the target output distribution unchanged under both greedy and sampled decoding. Code is available at this https URL.
Comments: 28 pages, 7 figures, 17 tables
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.00321 [cs.CL]
  (or arXiv:2610.00321v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.00321

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jungseob Lee [view email]
[v1] Tue, 29 Sep 2026 04:13:00 UTC (1,691 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org