跳到正文
arXiv:cs.LG· Manvi Jha, Zach Zhang, Zhichao Xu, Linbo Liu, Sai Muralidhar Jayanthi, Vinayak Arannil·· 4 小时前AI 评分37

APEX:更聪明地推测,而非更深

APEX: Speculate smarter, not deeper

AI 导读

APEX 是一种学习型控制器,通过请求级专家选择和块级深度自适应,在投机解码中平衡速度与草稿 token 浪费,并已集成进 vLLM。它用 Qwen3-8B 在六个工作负载上测试,最高实现 5.24X 加速;聚合评测中 APEX-S 达 4.27X,APEX-B 达 3.27X 且相比 k=16 固定 n-gram 投机将浪费 token 占比相对降低 41.0%。

正文

View PDF HTML (experimental)

Abstract:Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth. Fixed configurations cannot respond to changes in predictability, repetition, and acceptance during generation, so deeper drafting can increase wasted computation without proportional speedup. We introduce APEX, a learned controller that balances decoding speed and draft-token waste through request-level expert selection and block-level depth adaptation. APEX-Router selects among EAGLE-3, n-gram, and draft-model speculation for each request, while APEX-Depth adjusts draft length at each verification block using causal decoding signals and recent verifier feedback. APEX models accepted draft length as censored survival feedback, learning position-wise rejection hazards, block execution costs, and an action utility that balances throughput, accepted progress, and wasted tokens. This allows the controller to adapt speculation while retaining the target model's verification procedure. We integrate APEX into vLLM and evaluate it with Qwen3-8B across six workloads, achieving up to 5.24X speedup over autoregressive decoding. Across the aggregate evaluation, APEX-S achieves 4.27X speedup, while APEX-B achieves 3.27X speedup with a 41.0% relative reduction in wasted-token percentage compared with fixed n-gram speculation at k=16, providing distinct operating points for balancing acceleration and draft-token utilization.
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.07780 [cs.CL]
  (or arXiv:2610.07780v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.07780

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Manvi Jha [view email]
[v1] Tue, 6 Oct 2026 05:17:08 UTC (1,051 KB)

来源:arXiv:cs.LG · arxiv.org