跳到正文
arXiv:cs.LG· Andrew Cheng, Bobak T. Kiani, Yue M. Lu, Adityanarayanan Radhakrishnan·· 4 小时前AI 评分33

线性递归特征机(RFM)的精确动力学与有限样本轨迹恢复

Exact Dynamics and Finite-Sample Trajectory Recovery of Linear Recursive Feature Machines

AI 导读

研究将线性递归特征机(RFM)与迭代重加权最小二乘的已知联系,从插值设定扩展到带噪声的岭正则化多输出回归,其中数据为各向同性子高斯输入、目标由维度 d 的低秩教师矩阵生成。在 n 个样本下,学习到的特征矩阵以高概率以 O(√(d/n)) 的速率逼近其无限数据理想对应物。真实文本与单细胞基因表达数据实验展示了该线性模型学到的特征。

正文

View PDF HTML (experimental)

Abstract:Recursive feature machines (RFMs) learn representations of data by alternating between fitting a predictor to a dataset and updating features of that predictor using the average gradient outer product (AGOP). Connections between AGOPs and feature learning in neural networks motivate linear RFMs as a simple setting for analyzing how representations evolve during training. Here, we study the dynamics and statistics of linear RFM in noisy multi-output regression with isotropic sub-Gaussian input data and targets generated by a low-rank teacher matrix of dimension $d$. We extend the known connection between linear RFM and iteratively reweighted least squares from the interpolating setting to ridge-regularized multi-output regression with noise. We show that the learned feature matrix remains close to its infinite-data ideal counterpart at every iteration. Namely, for $n$ samples, we show the error in the feature matrix decays as $O(\sqrt{d/n})$ with high probability. Experiments on real-world text and single-cell gene-expression data illustrate the features learned by this simple linear model.
Comments: 51 pages, 8 figures
Subjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as: arXiv:2610.09196 [cs.LG]
  (or arXiv:2610.09196v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09196

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Andrew Cheng [view email]
[v1] Tue, 6 Oct 2026 22:45:58 UTC (516 KB)

来源:arXiv:cs.LG · arxiv.org