arXiv:cs.LG(机器学习,全量分类)· Haotian Gu, Yizhou Xu, Lenka Zdeborov\'a·· 14 小时前AI 评分28
单指标目标的上下文学习:核学习器与特征学习器对比
In-context Learning of Single-index Targets: Comparing Kernel and Feature Learners
AI 导读
一项理论研究对比了两种单层注意力架构在同一族单指标任务上的表现:核学习器先用固定非线性特征映射处理输入再做线性注意力,特征学习器则直接对原始输入做注意力并接可学习的非线性读出。作者用复本方法推导出两者在记忆与泛化误差上的预测,涵盖预训练规模、任务池多样性和训练/推理上下文长度的影响,预测与大量数值实验吻合,并给出相图刻画各架构何时占优,还发现两者在上下文长度缩放上存在定性差异。
正文
Abstract:In-context learning (ICL) enables a pretrained model to infer a task from demonstrations without updating its parameters. While much of the existing theory focuses on linear target functions, in this paper we study nonlinear cases by comparing two one-layer attention architectures on the same family of single-index tasks. A kernel learner first maps inputs through a fixed nonlinear feature map and then applies linear attention, whereas a feature learner applies attention to the original input, followed by a learned nonlinear readout. We derive predictions for their memorization and generalization errors using the replica method, retaining the effects of pretraining size, task-pool diversity, and training and inference context lengths. The resulting predictions closely match numerical experiments across a broad range of regimes. Our analysis yields phase diagrams that characterize when each architecture is advantageous as the amount of pretraining data, task diversity, and context lengths vary. We further identify qualitatively different context-length scalings for the two learners. Together, these results clarify how architectural choices interact with the dataset and govern nonlinear in-context learning.
| Subjects: | Machine Learning (cs.LG); Disordered Systems and Neural Networks (cond-mat.dis-nn); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.01712 [cs.LG] |
| (or arXiv:2610.01712v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01712 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yizhou Xu [view email]
[v1]
Thu, 1 Oct 2026 13:54:52 UTC (1,412 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org