跳到正文
arXiv:cs.CL· Vansh Bansal, Cholyeon Cho, Syamantak Kumar, Sujay Sanghavi, Purnamrita Sarkar·· 4 小时前AI 评分46

走快点但要小心:理解掩码扩散模型中的并行采样

Walk fast but be careful: Understanding Parallel Sampling in Masked Diffusion

AI 导读

研究以图上随机游走为可验证沙盒,分析掩码扩散模型(MDM)的并行采样策略,证明基于最低熵等分数的并行解掩码并不总优于随机并行采样,性能取决于图结构带来的条件依赖结构。

正文

View PDF HTML (experimental)

Abstract:In this paper, we use random walks on graphs as a verifiable sandbox for studying parallel sampling strategies in masked diffusion models (MDMs). We train an MDM on random walk samples from a fixed graph. The graph and transition kernel are never shown to the model and serve as latent structure that is both controllable and enables evaluation. The framework provides a validity check for generated walks and a measure of distributional fidelity through the estimated transition kernel. Using simple graphs, we theoretically prove that parallel unmasking via widely used scores such as lowest entropy is not uniformly better than random parallel sampling; even with exact conditional probabilities, performance critically depends on the conditional dependence structure induced by the graph, a phenomenon difficult to isolate in benchmarks like Sudoku. We also develop training-free bisection samplers for MDMs, which take logarithmically many steps in the sequence length and are provably exact for random walks if the learned marginals are exact. Experiments on graph-walk tasks confirm that different parallel samplers perform better on different graph structures. Experiments on pretrained MDMs show that bisection-style samplers provide strong speed-quality tradeoffs on OpenWebText generation and reasoning benchmarks including GSM8K, MBPP, and HumanEval. Together, these results use graph walks to uncover conditional dependence as a key principle of parallel MDM sampling and translate this insight into efficient samplers that transfer to language generation and reasoning.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2606.22976 [cs.LG]
  (or arXiv:2606.22976v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2606.22976

arXiv-issued DOI via DataCite

Submission history

From: Cholyeon Cho [view email]
[v1] Mon, 22 Jun 2026 07:56:50 UTC (2,272 KB)
[v2] Wed, 7 Oct 2026 14:59:29 UTC (691 KB)

来源:arXiv:cs.CL · arxiv.org