跳到正文
arXiv:cs.AI· Minghan Jiang, Jiayi Wang, Shuaiting Li, Haibin Shen, Kejie Huang·· 4 小时前

DynaTE:通过动态 Token 执行加速扩散 LLM

DynaTE: Accelerating Diffusion LLMs via Dynamic Token Execution

AI 导读

DynaTE 是一种面向扩散 LLM(dLLM)的软硬件协同设计架构,通过跳过低效用 token 计算、利用动态 token 依赖的 FLDD 机制减少去噪迭代次数,并用流式词表引擎处理不规则输出。

正文

View PDF HTML (experimental)

Abstract:Diffusion-based LLMs (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs by enabling bidirectional parallel refinement, alleviating the sequential decoding bottleneck of AR generation. However, their parallel iterative refinement mismatches AR accelerators optimized for sequential decoding and their discrete token generation differs from DiT accelerators designed for continuous denoising. Recent dLLM accelerators have explored workload-specific optimizations to reduce vocabulary processing overhead and redundant computation across denoising iterations. However, these approaches retain all tokens in parallel execution, despite varying token refinement utility and execution requirements.
This paper presents DynaTE, a hardware--software co-design architecture that dynamically adapts accelerator execution to evolving token states during dLLM decoding. DynaTE first enables adaptive token execution by skipping low-utility token computation, while a dimension-reconfigurable PE array maintains high utilization under varying active-token patterns. Second, DynaTE exploits dynamic token dependencies through FLDD to refine a small number of locally dependent tokens within the current iteration, reducing the overall number of denoising iterations, while a Merge--Split--Merge dataflow hides the resulting serial overhead. Third, a streaming vocabulary engine interleaves multiple token streams from the LM head to accommodate irregular output variations caused by selective token computation and uneven vocabulary-selection demands.
Evaluated on two representative dLLMs, DynaTE achieves 2.05--2.78$\times$ speedup and 2.99--3.93$\times$ higher energy efficiency over state-of-the-art dLLM accelerators, while delivering 2.55$\times$ speedup and 6.07$\times$ higher energy efficiency over Jetson AGX Orin.
Subjects: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11284 [cs.AR]
  (or arXiv:2610.11284v1 [cs.AR] for this version)
  https://doi.org/10.48550/arXiv.2610.11284

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Minghan Jiang [view email]
[v1] Thu, 8 Oct 2026 05:49:50 UTC (1,288 KB)

来源:arXiv:cs.AI · arxiv.org