跳到正文
arXiv:cs.LG· Saptarshi Chakraborty, Quentin Berthet, Peter L. Bartlett·· 3 小时前AI 评分38

流匹配模型对内在低维数据的泛化性质研究

Generalization Properties of Score-matching Diffusion Models for Intrinsically Low-dimensional Data

AI 导读

研究针对流匹配模型推导出有限样本泛化误差界,在 Wasserstein-p 距离下,学习分布与真实分布的误差以 n^{-1/d}+n^{-1/(2p)}(log(1/ξ))^{1/(2p)} 收敛,收敛指数取决于数据的内在维度而非环境维度。该结果在比现有分析更宽松的假设下成立,表明流匹配能自适应数据的内在几何结构并缓解维度灾难,为结构化数据分布上的实证成功提供理论解释。

正文

View PDF HTML (experimental)

Abstract:Despite the remarkable empirical success of flow-matching models, their statistical generalization guarantees remain underdeveloped. Existing analyses often impose restrictive assumptions on the estimated velocity field and yield convergence rates that fail to reflect the intrinsic low-dimensional structure common in real data, such as natural images and molecular geometries. In this work, we study the statistical generalization of flow-matching models for learning an unknown distribution $P_{\mathrm{data}}$ from finitely many samples. We derive finite-sample error bounds on the learned generative distribution, measured in the Wasserstein-$p$ distance, for all $p\geq 1$. Specifically, given $n$ i.i.d. samples from $P_{\mathrm{data}}$, we show that, for every $d>d_p^\ast(P_{\mathrm{data}})$ and appropriately chosen network architectures and hyperparameters, the learned distribution $\widehat{P}^{\mathrm{FM}}$ satisfies $ \mathbb{W}_p(\widehat{P}^{\mathrm{FM}},P_{\mathrm{data}}) \lesssim n^{-1/d}+n^{-1/(2p)}\bigl(\log(1/\xi)\bigr)^{1/(2p)}$ with probability at least $1-\xi$, where $d_p^\ast(P_{\mathrm{data}})$ denotes the Wasserstein-$p$ dimension of the target measure. Our results demonstrate that flow matching naturally adapts to the intrinsic geometry of data and mitigates the curse of dimensionality, as the convergence exponent depends on the intrinsic rather than ambient dimension. These guarantees remain meaningful in high-dimensional regimes and provide a theoretical explanation for the empirical success of flow matching on structured data distributions under substantially more relaxed assumptions than those in existing analyses.
Subjects: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Statistics Theory (math.ST)
Cite as: arXiv:2610.02663 [stat.ML]
  (or arXiv:2610.02663v1 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2610.02663

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Saptarshi Chakraborty [view email]
[v1] Fri, 2 Oct 2026 01:32:38 UTC (64 KB)

来源:arXiv:cs.LG · arxiv.org