arXiv:cs.LG· Stanislas Strasman (SU, LPSM), Sobihan Surendran (SU, LPSM), Sylvain Le Corff (SU, LPSM)·· 6 小时前AI 评分31
随机梯度下降在基于分数的生成模型中的非渐近收敛性
Non-asymptotic Convergence of Stochastic Gradient Descent in Score-based Generative Models
AI 导读
研究针对基于分数的生成模型(SGMs)训练中随机梯度下降(SGD)的优化动力学展开分析。对一般分数参数化,给出了加权去噪分数匹配目标下SGD的非凸分析,明确了优化界与损失权重及时间采样分布的关系;对过参数化两层ReLU网络,通过Neural Tangent Kernel分析推导出沿SGD轨迹的分数近似误差界。
正文
Abstract:Score-based Generative Models (SGMs) have achieved impressive performance in data generation across a wide range of applications. While the statistical properties of their sampling procedures are increasingly well understood, the optimization dynamics underlying their training remain less explored. SGMs are typically trained by minimizing a weighted denoising score-matching objective, yet optimization guarantees with stochastic gradients remain limited. In this work, we study Stochastic Gradient Descent (SGD) for SGMs, contributing results in two complementary regimes. For general score parameterizations, we derive a non-convex analysis of SGD for the weighted denoising score-matching objective, making explicit how the resulting optimization bound depends on the loss weighting and time-sampling distribution. We then consider overparameterized two-layer ReLU networks and develop a Neural Tangent Kernel analysis tailored to diffusion training with stochastic gradients, yielding score-approximation error bounds along the SGD trajectory. Our analysis quantifies the role of the reweighting factor in these bounds, providing a theoretical characterization of weighting choices used in practice.
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.04775 [stat.ML] |
| (or arXiv:2607.04775v2 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2607.04775 arXiv-issued DOI via DataCite |
Submission history
From: Stanislas Strasman [view email] [via CCSD proxy]
[v1]
Mon, 6 Jul 2026 08:07:15 UTC (57 KB)
[v2]
Wed, 7 Oct 2026 13:42:39 UTC (73 KB)
来源:arXiv:cs.LG · arxiv.org