跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Yair Schiff, Omer Belhasin, Roy Uziel, Matan Rusanovsky, Ran Zilberstein, Marianne Arriola, Gilad Turok, Guanghan Wang, Volodymyr Kuleshov, Michael Elad·· 14 小时前AI 评分43

Clock Diffusion:高效半自回归连续扩散语言模型

Clock Diffusion: Efficient Semi-Autoregressive Continuous Diffusion Language Models

AI 导读

研究者提出 Clock Diffusion 框架,通过位置相关噪声调度实现半自回归连续扩散语言模型,支持变长生成和 KV cache,并给出 block 与滑动窗口两种生成方式。

正文

View PDF HTML (experimental)

Abstract:Recent works on continuous diffusion for discrete data have demonstrated performance on par with comparable discrete diffusion models. However, these continuous counterparts lack key features that are essential to practical use as language models, namely variable-length generation and support for a key-value cache, and they still lag behind the frontier of autoregressive and discrete diffusion quality. In this work, we address these limitations. We do so by introducing a model parameterization that uses position-dependent noise schedules to define semi-autoregressive (SAR) continuous diffusion language models (DLMs). Together with efficient training and sampling algorithms, we call this framework Clock Diffusion, and we present two special cases of our method: block and sliding window generation. We then define ClockDLMs, a family of Gaussian DLMs based on sliding window Clock Diffusion that attain state-of-the-art diffusion likelihood bounds on OpenWebText, even beating the performant block SAR discrete diffusion models. ClockDLMs trained on TinyGSM also substantially outperform continuous baselines on the GSM8K benchmark and match and exceed comparable SAR discrete diffusion models. Finally, building on our parameterization, we propose more efficient samplers that we dub Cache Grab, which adapt techniques from accelerated inference in discrete diffusion, such as committing tokens whose probabilities exceed a confidence threshold and self-speculative decoding, further improving our models' quality and efficiency.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.00894 [cs.LG]
  (or arXiv:2610.00894v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00894

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yair Schiff [view email]
[v1] Thu, 1 Oct 2026 01:17:31 UTC (197 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org