跳到正文
arXiv:cs.AI· Arnold Caleb Asiimwe, William Yang, Sanghyuk Chun, Esin Tureci, Olga Russakovsky·· 5 小时前AI 评分44

一步生成模型中的「深度即时间」现象:扩散去噪轨迹如何重组进网络深度

Depth as Time in One-Step Generative Models

AI 导读

研究者提出「深度即时间」经验观察:多步扩散的去噪计算在一步生成模型中被重组到单次前向传播的网络深度上,可用模型自身输出头解码中间层恢复。该现象依赖 flow map 所训练的传输任务,MeanFlow 在更短传输区间内单次评估同时呈现去噪与再加噪,而 drifting models 等无时间索引传输任务的生成器不具备此特性。

正文

View PDF HTML (experimental)

Abstract:The recent wave of one-step generative models, which compress the multi-step trajectory of diffusion via either distillation or learned flow maps, has reached an inflection point where they can generate high-quality images. Here, we ask a natural question that follows from these advances: what happens to the denoising trajectory of multi-step diffusion when generation is compressed into a single forward pass? We offer an empirical observation we call \textit{depth as time}: the denoising computation that multi-step diffusion performs across sampling steps appears to unfold across the depth of a single forward pass, and can be recovered by decoding intermediate layers with the model's own output head. Most interestingly, we show that this depthwise computation depends on the transport task a flow map is trained to solve. The most surprising case is MeanFlow, where probing shorter transport intervals reveals both denoising and renoising within a single network evaluation. In contrast, generators trained without a time-indexed transport task, such as drifting models, do not exhibit the same depthwise denoising. Consequently, we show that models that exhibit the depthwise denoising phenomenon are more compressible across the layerwise computation: a MeanFlow \texttt{SiT-L/2} model can be compressed by $16.6\times$ in parameters into a single time-conditioned block. We offer an explanation for this denoise-then-renoise behavior and show that, when we treat the layerwise computation explicitly as a flow, a single time-conditioned block can be trained to denoise across layers, compressing a MeanFlow \texttt{SiT-L/2} model by $16.6\times$ in parameters. Together, these results suggest that the temporal computation of diffusion is not eliminated by one-step generation, but reorganized across network depth.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03626 [cs.AI]
  (or arXiv:2610.03626v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.03626

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Arnold Caleb Asiimwe [view email]
[v1] Fri, 2 Oct 2026 17:20:08 UTC (36,032 KB)

来源:arXiv:cs.AI · arxiv.org