跳到正文
arXiv:cs.LG· Nikhil Verma, Siddharthan Dileep, Anoop Singh, Srikanth Sastry, Ramya Hebbalaguppe, Sayan Ranu, N. M. Anoop Krishnan·· 4 小时前

扩散模型的记忆化早期信号:基于盆地几何与循环去噪

Early Signatures of Memorization in Diffusion Models via Basin Geometry and Cyclic Denoising

AI 导读

研究发现扩散模型的记忆化会在生成样本出现前先编码进学习到的能量景观几何中,这种状态称为潜在记忆化。通过 score divergence 和盆地体积,可在首个记忆样本出现前观察到训练样本周围的局部盆地,其起始遵循与记忆化时间相同的 O(n) 缩放。

正文

View PDF HTML (experimental)

Abstract:Diffusion models generalize early in training and later reproduce individual training samples. Standard tests detect memorization only once one-shot generation produces near-copies, leaving a released model unaudited until its outputs fail. We show that memorization is encoded in the geometry of the learned energy landscape before it appears in generated samples, a state we call latent memorization. Using score divergence and basin volume, we find that localized basins form around training samples and separate them from held-out samples before the first memorized sample appears, with an onset that follows the same $O(n)$ scaling as the memorization time. We probe these basins with cyclic denoising, which repeatedly applies partial noising and denoising. Under the exact empirical score, we prove that cycling started near an isolated training sample recovers it and returns to it over any finite number of cycles with high probability. In trained models, cycling recovers training images from CelebA and CIFAR-10 checkpoints whose one-shot samples contain no copies, and at a CelebA checkpoint with 0.1% one-shot copies, 500 cycles raise the memorized fraction above 30%. Cycling also reveals degenerate attractors that match no single training image and fade as training proceeds, so residence in a basin does not by itself imply memorization. These findings hold on a Gaussian mixture, CelebA, and CIFAR-10 across optimizers, architectures, noise schedules, and training-set sizes, and extend to off-the-shelf Stable Diffusion v1.4, where the cycled conditional-unconditional divergence gap separates memorized from non-memorized prompts with an AUC of 0.944 and a TPR of 0.866 at 1% FPR. More broadly, what a diffusion model has memorized is a property of the geometry and stability of its learned distribution, and assessing it requires examining this structure rather than generated outputs alone.
Comments: 42 pages, 24 figures, 7 tables
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.11670 [cs.LG]
  (or arXiv:2610.11670v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.11670

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Nikhil Verma [view email]
[v1] Thu, 8 Oct 2026 10:47:19 UTC (12,812 KB)

来源:arXiv:cs.LG · arxiv.org