跳到正文
arXiv:cs.LG· Jivan Waber, Vanessa Piccolo, Yatin Dandi, Florent Krzakala·· 4 小时前AI 评分35

对称感知特征学习的多项式分离:多指标模型研究

Symmetry-Aware Feature Learning: A Polynomial Separation for Multi-Index Models

AI 导读

研究证明对称感知与对称无关特征学习间存在多项式样本复杂度分离。在$\mathbb{R}^d$上秩$r=\Theta(d^\delta)$的多指标模型中,权重共享网络与全群数据增强在$\widetilde{\Theta}(d^{p-1})$样本内实现弱方向恢复,而对称无关学习需$\widetilde{\Theta}(rd^{p-1})$样本。

正文

View PDF HTML (experimental)

Abstract:We establish a polynomial sample complexity separation between symmetry-aware and symmetry-agnostic feature learning. We study growing-rank multi-index models with high-dimensional Gaussian covariates in $\mathbb{R}^d$ and $r=\Theta(d^\delta)$ teacher directions forming a cyclic symmetry orbit, where $0<\delta<1/2$. We compare three ways of exploiting this structure: architectural weight sharing, data augmentation over the full symmetry group, and learning without access to the symmetry. In particular, we analyze a symmetry-tied convolutional network, an untied network, and the same untied network trained with full-group data augmentation, using spherical online SGD with correlation loss. For a class of polynomial links with information exponent $p\ge3$, we prove matching sample complexity bounds up to logarithmic factors: the tied and augmented learners achieve weak directional recovery in $\widetilde{\Theta}(d^{p-1})$ samples, whereas the symmetry-agnostic learner requires $\widetilde{\Theta}(rd^{p-1})$. For the pure quadratic Hermite link, the same separation holds for weak recovery of the teacher subspace, with sample complexities $\widetilde{\Theta}(d)$ and $\widetilde{\Theta}(rd)$, respectively. Thus, full-group data augmentation matches the sample efficiency of architectural weight sharing, and both provide a polynomial advantage over training without symmetry. For $p\ge3$, the proof reveals a two-stage mechanism: fluctuations at initialization select one direction in the teacher orbit, after which localized growth amplifies its overlap to the weak recovery scale while competing overlaps remain near their initialization scale.
Comments: 71 pages, 3 figures
Subjects: Machine Learning (cs.LG); Probability (math.PR); Machine Learning (stat.ML)
Cite as: arXiv:2610.08420 [cs.LG]
  (or arXiv:2610.08420v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.08420

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Vanessa Piccolo [view email]
[v1] Tue, 6 Oct 2026 14:23:00 UTC (195 KB)

来源:arXiv:cs.LG · arxiv.org