arXiv:cs.LG· Chung-En Ho, Weiyu Sun, Cheng-Jhih Shih, He Li, Yong Liu, Yingyan Celine Lin·· 7 小时前AI 评分39
SpecFold:折叠多分支冗余加速扩散语言模型投机解码
SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models
AI 导读
SpecFold 是一种算法-系统协同设计,通过折叠注意力与 FFN 选择性复用父分支计算,消除扩散语言模型多分支投机验证中的冗余,实现最高 1.64x(对比 Spiffy)和 1.99x(对比原始解码)的吞吐提升。该方法在五款模型、五个标准基准上保持任务性能相当,可与现有 DLLM 投机策略及时间缓存正交组合。
正文
Abstract:Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit temporal redundancy across denoising steps, we identify a complementary redundancy axis within each speculative verification step: multi-branch computational redundancy. During speculative verification, draft branches inherit most tokens from their parents while unmasking a small set of additional positions, causing large portions of hidden states to remain highly similar across branches. We propose SpecFold, an algorithm-system co-design that exploits this multi-branch redundancy to reduce the cost of multi-branch speculative verification. Algorithmically, SpecFold performs token-level residual gating and selectively reuses parent computation through folded attention and FFN while preserving residual hidden states. Systemically, a Triton kernel implementation translates this fine-grained reuse into end-to-end throughput gains through efficient sparse multi-branch execution. SpecFold is orthogonal to temporal caching and compatible with existing DLLM speculation strategies. Across two DLLM families, five models, and five standard benchmarks, SpecFold achieves up to 1.64x throughput over Spiffy and up to 1.99x over vanilla decoding, while maintaining comparable task performance.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.04875 [cs.AI] |
| (or arXiv:2610.04875v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.04875 arXiv-issued DOI via DataCite |
Submission history
From: Chung-En Ho [view email]
[v1]
Sun, 4 Oct 2026 02:25:38 UTC (453 KB)
[v2]
Tue, 6 Oct 2026 03:42:24 UTC (453 KB)
来源:arXiv:cs.LG · arxiv.org