跳到正文
arXiv:cs.LG· Pawe{\l} Skier\'s, Emilia Kaczmarczyk, Tomasz Trzci\'nski, Kamil Deja·· 4 小时前AI 评分44

ELROND:探索并分解扩散模型的内在能力

ELROND: Exploring and decomposing intrinsic capabilities of diffusion models

AI 导读

ELROND 通过反向传播固定提示词不同随机生成结果之间的差异来收集梯度,并用主成分分析或稀疏自编码器将其分解为可解释方向,从而在给定条件下恢复扩散模型生成流形的切空间。该方法在受控环境中准确恢复该子空间,并在大规模模型上通过行为验证,其恢复的结构比依赖外部表示的方法能更广地探索模型能力,且无需重训练即可缓解蒸馏模型的模式崩溃。

正文

View PDF HTML (experimental)

Abstract:A single text prompt passed to a diffusion model yields a wide range of visual outputs determined solely by a stochastic process, leaving users with no direct control over which semantic variations appear. Exploring this range is difficult: random search offers no guarantee of covering it, while prompt editing is coarse, as even a small change in wording can substantially alter the generated image. We argue that systematic exploration instead requires recovering how the model itself organizes the conditional distributions it can produce. We formalize this structure as a generative manifold, and present ELROND, a method for recovering its tangent space at a given conditioning. To that end, we collect gradients obtained by backpropagating the differences between stochastic realizations of a fixed prompt, and decompose them into interpretable directions using Principal Component Analysis or a Sparse Autoencoder. We show that our method recovers this subspace accurately in a controlled setting and validate it behaviorally on large-scale models. We also demonstrate that recovered structure enables broader exploration of the model's capabilities than methods relying on external representations, and mitigates mode collapse in distilled models without retraining.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2602.10216 [cs.LG]
  (or arXiv:2602.10216v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2602.10216

arXiv-issued DOI via DataCite

Submission history

From: Paweł Skierś [view email]
[v1] Tue, 10 Feb 2026 19:07:15 UTC (15,557 KB)
[v2] Wed, 7 Oct 2026 11:42:21 UTC (24,852 KB)

来源:arXiv:cs.LG · arxiv.org