跳到正文
arXiv:cs.LG· Mary Letey, Arman Rysmakhanov, Yue M. Lu, Cengiz Pehlevan, Jacob Zavatone-Veth·· 3 小时前

线性注意力能从非线性教师那里学到什么?上下文学习理论扩展至非线性单指标目标

What can linear attention learn from nonlinear teachers in-context?

AI 导读

一项理论研究将线性注意力的上下文学习理论扩展到非线性单指标目标 y=f(x^⊤w)+ε,提出"非线性-噪声等价"结论:线性注意力只提取 f 的线性 Hermite 分量,其余非线性结构作为有效噪声计入泛化误差。该结论使线性理论结果可迁移至非线性任务,并揭示任务多样性增加时从任务记忆到任务泛化的转变及有限预训练数据的影响,同时指出简化线性注意力模型的局限。

正文

View PDF HTML (experimental)

Abstract:Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers. For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour. We extend this theory to nonlinear single-index targets, $y=f(x^\top w)+\varepsilon $. Our main result establishes a nonlinearity-noise equivalence: linear attention extracts only the linear Hermite component of $f$, while the remaining nonlinear structure contributes to the generalisation error as effective noise. This reduction allows results from the corresponding linear theory to be transferred to nonlinear tasks. We illustrate its implications for finite pretraining data and for the transition from task memorisation to task generalisation as task diversity increases. These results identify a limitation of the reduced linear-attention model and provide a tractable starting point for studying nonlinear in-context learning.
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG)
Cite as: arXiv:2610.10761 [stat.ML]
  (or arXiv:2610.10761v1 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2610.10761

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mary Letey [view email]
[v1] Wed, 7 Oct 2026 18:24:13 UTC (62 KB)

来源:arXiv:cs.LG · arxiv.org