HC-DLM:分层连续扩散语言模型
Hierarchical Continuous Diffusion Language Models
研究者提出分层连续扩散语言模型(HC-DLM),将离散 token 生成与连续隐变量轨迹耦合进单一去噪过程,训练目标由 token 似然的变分下界推导而来。与把连续上下文附加到自足离散链上的近期方法不同,HC-DLM 让隐变量成为唯一持久生成状态,每步从中读出 token 并反馈为下一步隐变量更新的脚手架。
Published on Oct 1
Authors:
,
,
,
Abstract
Discrete diffusion language models offer a compelling alternative to autoregressive generation for tasks demanding bidirectional reasoning and global constraint satisfaction. Yet they share a structural bottleneck: when decoding in parallel, each token is sampled independently from its marginal, severing the statistical dependencies among the tokens decoded together. Continuous diffusion language models avoid this by denoising a shared continuous state, but their denoiser sees only that state, so nothing ties it to a valid token configuration until it is finally decoded. To address this, we propose Hierarchical Continuous Diffusion Language Models (HC-DLM), which couple discrete token generation with a continuous latent trajectory in a single, principled denoising process, whose training objective is derived from a variational bound on the token likelihood. In contrast to recent methods that attach continuous context to a self-contained discrete chain, HC-DLM makes the latent the only persistent generative state: tokens are read out from it at every step and feed back as a scaffold for the next latent update. On structured reasoning (Sudoku), mathematical planning (Countdown) and language modeling (LM1B), HC-DLM improves over discrete and continuous diffusion baselines at matched model size, in puzzle accuracy on Sudoku and Countdown and in generative perplexity on LM1B. Project page: https://hc-dlm.github.io/.
View arXiv page View PDF Project page GitHub 1 Add to collection
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2610.02193 in a model README.md to link it from this page.
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2610.02193 in a dataset README.md to link it from this page.
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2610.02193 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.
来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co