arXiv:cs.LG· Andrew Cheng, Bobak T. Kiani, Yue M. Lu, Adityanarayanan Radhakrishnan·· 4 小时前AI 评分33
线性递归特征机(RFM)的精确动力学与有限样本轨迹恢复
Exact Dynamics and Finite-Sample Trajectory Recovery of Linear Recursive Feature Machines
AI 导读
研究将线性递归特征机(RFM)与迭代重加权最小二乘的已知联系,从插值设定扩展到带噪声的岭正则化多输出回归,其中数据为各向同性子高斯输入、目标由维度 d 的低秩教师矩阵生成。在 n 个样本下,学习到的特征矩阵以高概率以 O(√(d/n)) 的速率逼近其无限数据理想对应物。真实文本与单细胞基因表达数据实验展示了该线性模型学到的特征。
正文
Abstract:Recursive feature machines (RFMs) learn representations of data by alternating between fitting a predictor to a dataset and updating features of that predictor using the average gradient outer product (AGOP). Connections between AGOPs and feature learning in neural networks motivate linear RFMs as a simple setting for analyzing how representations evolve during training. Here, we study the dynamics and statistics of linear RFM in noisy multi-output regression with isotropic sub-Gaussian input data and targets generated by a low-rank teacher matrix of dimension $d$. We extend the known connection between linear RFM and iteratively reweighted least squares from the interpolating setting to ridge-regularized multi-output regression with noise. We show that the learned feature matrix remains close to its infinite-data ideal counterpart at every iteration. Namely, for $n$ samples, we show the error in the feature matrix decays as $O(\sqrt{d/n})$ with high probability. Experiments on real-world text and single-cell gene-expression data illustrate the features learned by this simple linear model.
| Comments: | 51 pages, 8 figures |
| Subjects: | Machine Learning (cs.LG); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.09196 [cs.LG] |
| (or arXiv:2610.09196v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09196 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Andrew Cheng [view email]
[v1]
Tue, 6 Oct 2026 22:45:58 UTC (516 KB)
来源:arXiv:cs.LG · arxiv.org