arXiv:cs.LG· Xinsong Feng, Peng Du, Zhizhuo Yang, Daniel M. Bikel, Jiayun Wang, Haipeng Chen·· 4 小时前AI 评分41
BLD:将 1024 token 压缩为 64 个块隐变量的高效连续扩散模型
Denoising Blocks, Not Tokens: Efficient Compressed Continuous Diffusion with Branching Token Realization
AI 导读
研究者提出 Branching Latent Diffusion(BLD),将 1024 token 序列压缩为仅 64 个块隐变量(16 倍压缩),并通过分支 token 实现机制用并行局部 AR 分支解码每个隐变量。
正文
Abstract:Diffusion language models (DLMs) generate text through iterative parallel refinement, offering the potential for higher throughput than autoregressive (AR) decoding. However, most DLMs still maintain one generative state per token, so every denoising step processes a state sequence as long as the output sequence, limiting the throughput gains from parallel generation. Continuous DLMs provide an additional degree of freedom: a single continuous state can represent multiple tokens, allowing diffusion to operate on a much shorter latent sequence. We introduce \emph{Branching Latent Diffusion (BLD)}, which exploits this flexibility by compressing a 1024-token sequence into only 64 block latents, a $16\times$ reduction. BLD combines latent compression with \emph{branching token realization}, where each latent is decoded by a local AR branch and all branches run in parallel. Because strong compression makes joint latent generation difficult, BLD generates the latents in groups, conditioning each group on previously generated latents. In end-to-end evaluation on the same GPU, BLD reduces generation FLOPs by more than $80\times$ and increases throughput by more than $6\times$ relative to the similarly sized ELF-L baseline. Compared with the AR baseline, BLD achieves more than $6\times$ higher throughput and more than $4\times$ lower latency. Despite the compression, BLD maintains competitive local fluency and diversity, although long-range coherence remains challenging. Overall, BLD shows that moving diffusion from token-level states to compressed latent sequences can substantially improve the efficiency of long-sequence generation.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.09311 [cs.LG] |
| (or arXiv:2610.09311v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09311 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xinsong Feng [view email]
[v1]
Wed, 7 Oct 2026 02:08:22 UTC (310 KB)
来源:arXiv:cs.LG · arxiv.org