跳到正文
arXiv:cs.LG· Calvin Yeung, Prathyush Poduval, Ali Zakeri, Zhuowen Zou, Mohsen Imani·· 6 小时前AI 评分30

用于扩散模型可解释性的残差化时序稀疏自编码器 ReSAE

Residualized Temporal Sparse Autoencoders for Interpreting Diffusion Models

AI 导读

研究者提出残差化时序稀疏自编码器(ReSAE),通过拟合相邻时间步之间的线性预测器,用初始激活和线性动态无法解释的残差来表示扩散模型的激活轨迹。该方法等价于在线性动态诱导的度量下对原始轨迹训练 SAE。在 Stable Diffusion 1.5 和 Diffusion Transformer 上,ReSAE 特征覆盖整条去噪轨迹并能定位变化进入的时间点,可用于研究扩散模型如何随时间生成图像。

正文

View PDF HTML (experimental)

Abstract:Text-to-image diffusion models generate images by iterative denoising, so their internal layers produce trajectories of activations rather than single static representations. Sparse autoencoders (SAEs) have recently been used to decompose diffusion activations into interpretable features, but most approaches analyze individual timesteps or condition on time rather than learning from full trajectories. Training one SAE on whole trajectories would make each feature a single trajectory across timesteps, but adjacent activations are largely linearly predictable from one another, so such an SAE spends its latents on content carried forward from step to step. We introduce residualized temporal SAEs (ReSAE), which fit linear predictors between neighboring timesteps and represent each trajectory by its initial activation and the residuals these dynamics leave unexplained. Training an SAE on this representation is equivalent to training it on raw trajectories under a metric induced by the linear dynamics, and it encourages latents to capture structure beyond what is linearly predictable. Each latent's decoder direction maps back to activation space as a feature trajectory over denoising time. Across Stable Diffusion~1.5 and a Diffusion Transformer, ReSAE features span the whole trajectory while pinpointing when changes enter it, making ReSAE a natural tool for studying how a diffusion model generates an image over time.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2605.27813 [cs.CV]
  (or arXiv:2605.27813v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2605.27813

arXiv-issued DOI via DataCite

Submission history

From: Calvin Yeung [view email]
[v1] Wed, 27 May 2026 01:08:29 UTC (13,377 KB)
[v2] Wed, 7 Oct 2026 04:52:56 UTC (3,291 KB)

来源:arXiv:cs.LG · arxiv.org