跳到正文
arXiv:cs.LG· Yudou Tian, Neeraj Mohan Sushma, Harshvardhan Mestha, Nicolo Colombo, David Kappel, Anand Subramoney·· 4 小时前AI 评分43

对角线性循环网络 GRIL 如何在循环状态中实现上下文梯度下降

Learning in the Recurrent State: Gradient Descent with Linear Recurrent Networks

AI 导读

研究者提出 GRIL(Gradient-based Recurrent In-context Learner),一种对角线性循环网络,将监督梯度步分解为短窗口互积写入与对下一查询的乘法读出。

正文

View PDF HTML (experimental)

Abstract:In-context learning lets a sequence model adapt to a new task from examples in its input. A prominent line of work shows how self-attention can be constructed to implement gradient descent on a linear predictor fit to the in-context examples during the forward pass. State-space models (SSMs) and other linear recurrent networks (LRNNs) model sequences at linear time cost, but it is unclear how their recurrent update could carry out the same in-context gradient descent. We introduce Gradient-based Recurrent In-context Learner (GRIL), a diagonal LRNN that factorizes a supervised gradient step into a short-window cross-product write and a multiplicative readout of the next query. For linear regression, this construction accumulates the context gradient in a matrix state and applies it in a single forward pass, with $O(f^2)$ learned degrees of freedom. The same design extends to multi-step updates and cross-entropy classification, with a limited MLP-based extension to non-linear regression. We show empirically that trained GRILs recover the behavior and parameters analytically predicted by the construction on synthetic ICL tasks. Furthermore, the same architecture can be extended and trained on general-purpose benchmarks, including Long Range Arena, language modeling and associative recall. Together, these results establish windowed cross-product self-attention as a concrete inductive bias that lets LRNNs learn in context through gradient-descent-like updates, while remaining trainable on general-purpose tasks.
Comments: 40 pages, 10 figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Neural and Evolutionary Computing (cs.NE)
Cite as: arXiv:2410.11687 [cs.LG]
  (or arXiv:2410.11687v4 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2410.11687

arXiv-issued DOI via DataCite

Submission history

From: Yudou Tian [view email]
[v1] Tue, 15 Oct 2024 15:22:38 UTC (585 KB)
[v2] Tue, 18 Feb 2025 18:55:39 UTC (388 KB)
[v3] Mon, 15 Jun 2026 15:20:33 UTC (776 KB)
[v4] Wed, 7 Oct 2026 12:40:39 UTC (797 KB)

来源:arXiv:cs.LG · arxiv.org