跳到正文
arXiv:cs.LG· Ruoyu Zhao, Yuting Chen, Jinheng Zhang, Zhehao Zou, Tong Che·· 3 小时前AI 评分36

共享高斯化:高斯正则项能为对比学习证明什么,又遗漏了什么

Shared Gaussianization: What Gaussian Regularizers Certify About Contrastive Learning, and What They Miss

AI 导读

研究提出共享高斯化(SG),一种对两个归一化视图均值做特征函数高斯性检验的方法,因视图不一致会缩短均值,单一检验即可同时检测错位与非均匀性。SG 在对齐且均匀的 InfoNCE 总体最优解处恰好为零,并给出以 SG 损失平方根为上界的 InfoNCE 超额界,该平方根速率与无穷维常数均为紧的。

正文

View PDF HTML (experimental)

Abstract:What can a distribution-matching regularizer such as SIGReg in LeJEPA certify about contrastive learning? We study shared Gaussianization (SG), a characteristic-function Gaussianity test on the average of two normalized views, scaled by an independent $\chi_d$ radius. Because disagreeing views shorten the average, one test detects both misalignment and non-uniformity. SG vanishes exactly at the aligned, uniform minimizers of population InfoNCE, and under equal marginals it bounds the InfoNCE excess by $4\cdot 3^{3/4}\beta$ times the square root of the SG loss, plus a term linear in the loss. The square-root rate and this dimension-free constant are sharp, and no squared mean-embedding distance on view pairs achieves a faster rate. With an explicit alignment term, a rotation-invariant uniformity test gives a linear bound if and only if its spectrum dominates that of InfoNCE's kernel $e^{\beta u^\top v}$; SG's own test does, Gaussian kernels $e^{-\gamma \|u-v\|^2}$ qualify exactly when $\gamma \ge \beta/2$, and moment matching never does. Away from the optimum, the objectives differ. Along an isotropic nuisance channel, pure SG lowers its loss by adding per-view nuisance whenever the shared code is non-uniform. An alignment weight above the channel's gain makes the nuisance-free solution a strict local minimizer; for LeJEPA, the same rule gives a critical SIGReg weight that decreases with the batch size. At finite batch size, an off-diagonal U-statistic removes a plug-in bias toward misalignment. In controlled latent-variable models, pure SG retains per-view style, an alignment weight above the measured gain removes it, and for LeJEPA at three batch sizes the measured gain separates the encoders that retain style from those that do not. InfoNCE training also reaches a lower SG$_{0.2}$ loss than SG$_{0.2}$ training from scratch, which points to an optimization gap.
Comments: 27 pages, 4 figures, 3 tables. Ruoyu Zhao and Yuting Chen contributed equally; Tong Che is the project lead
Subjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as: arXiv:2610.10299 [cs.LG]
  (or arXiv:2610.10299v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.10299

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ruoyu Zhao [view email]
[v1] Wed, 7 Oct 2026 15:53:56 UTC (149 KB)

来源:arXiv:cs.LG · arxiv.org