跳到正文
arXiv:cs.LG· Evan Dogariu, Joan Bruna·· 3 小时前

Barron 最优传输 I:生成建模

Barron Optimal Transport I: Generative Modeling

AI 导读

研究提出以 Barron 能量替代经典最优传输中的平均动能 L² 能量,构建概率测度空间上的新度量,用于衡量神经网络的表示复杂度。基于该度量,作者证明了由神经网络前推高斯生成的数据存在超多项式 score 逼近下界,从而量化了扩散生成建模在 Barron 几何下的次优性。

正文

View PDF HTML (experimental)

Abstract:Motivated by recent applications in generative modeling and sampling, we introduce a framework for optimal measure transport where cost captures the notion of neural network complexity. In transport-based generative models, samples from a reference distribution (e.g. Gaussian) are mapped to samples of a target distribution along ordinary or stochastic differential equations. These are implemented as deep residual networks when discretized in time, where each hidden layer approximates the associated instantaneous velocity. Thus, given a pair of target and reference measures, a natural question is to search for the most efficient neural representation that implements this transport.
Our starting point is the kinetic formulation of OT, due to Benamou and Brenier. We replace the average kinetic $L^2$ energy by the \emph{Barron} energy \cite{bach2017breaking, ma2022barron}, a natural norm which measures the complexity of representing a given vector field with a neural hidden layer, and which captures the adaptive properties of feature learning. This defines a metric on the space of probability measures, complementing existing Wasserstein and Stein geometries.
In this work we examine the properties of this metric in the context of generative modeling. As a first application, we quantify the suboptimality of diffusion generative modeling in the Barron geometry by establishing super-polynomial score approximation lower bounds for data generated by neural network pushforwards of the Gaussian. We then investigate the benefit of adaptivity as a way to study alternative generative models. In a companion paper \cite{companionpaper} we leverage the Barron transport geometry for sampling applications, extending the scope of Stein variational gradient methods via feature adaptation.
Subjects: Machine Learning (cs.LG); Analysis of PDEs (math.AP)
Cite as: arXiv:2610.10875 [cs.LG]
  (or arXiv:2610.10875v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.10875

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Joan Bruna [view email]
[v1] Wed, 7 Oct 2026 20:22:00 UTC (452 KB)

来源:arXiv:cs.LG · arxiv.org