arXiv:cs.LG· Gwangho Kim, Sungyoon Lee·· 3 小时前
条件扩散模型中的记忆与恶性泛化:随机特征理论分析
Memorization and Malign Generalization in Conditional Diffusion Models with Random Features
AI 导读
研究者通过分析高维比例极限下的随机特征条件得分模型,推导出训练与测试损失的渐近表达式,并据此提出"恶性泛化"现象:在过参数化区间,增大模型宽度能改善条件依赖均值的预测,同时降低条件内预测方差。研究还发现条件信息量越大,训练样本会在更小的宽度下被记忆。U-Net 架构在真实数据上的实验支持了这些理论结论。
正文
Abstract:Conditional diffusion models generate diverse, novel, and high-quality samples under prescribed conditions. However, theoretical understanding of their memorization and generalization remains limited, while recent works have characterized these behaviors primarily in unconditional settings. In this work, we analyze a random-feature conditional score model in the high-dimensional proportional limit, deriving asymptotic expressions for training and test losses. By decomposing the test loss, we show that in the overparameterized regime, increasing model width improves prediction of the condition-dependent mean while reducing within-condition prediction variance, a phenomenon we term "malign generalization." Furthermore, analyzing the training loss reveals that more informative conditions lead to memorization of training samples at smaller widths. These theoretical findings are supported by experiments with U-Net architectures on realistic data.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.11288 [cs.LG] |
| (or arXiv:2610.11288v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11288 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Gwangho Kim [view email]
[v1]
Thu, 8 Oct 2026 05:52:04 UTC (4,168 KB)
来源:arXiv:cs.LG · arxiv.org