跳到正文
arXiv:cs.LG· Giovanni Monea, Keshav Ramji, Yousef El-Kurdi, Luis A. Lastras, Yoav Artzi, Nathan Godey, Ram\'on Fernandez Astudillo·· 3 小时前AI 评分39

共享内存为何在循环 Transformer 中意外有效

The Surprising Effectiveness of Shared Memory in Looped Transformers

AI 导读

研究者预训练循环语言模型共享内存:仅首次递归写入 key-value cache,后续递归读取该缓存并保留一小段自身窗口。

正文

View PDF HTML (experimental)

Abstract:Looped Transformers apply the same layers several times per token, adding compute to improve quality without more parameters. Each recursion, however, writes its own key-value cache, so memory still grows with compute. Inference-time techniques can shrink this cache at a cost in quality. We pretrain looped language models to share memory: only the first recursion writes a cache, and later recursions read it while keeping a short window of their own. Surprisingly, we find that sharing memory does not cost quality and instead improves it. At 150M-1B parameters, our Looped Prediction Transformer (LPT) and its hybrid variant set a new quality-memory frontier for looped models: with five recursions, the hybrid lowers validation perplexity on FineWeb-Edu by 1.12-1.82 relative to a same-size standard Transformer while using 76-79% less context memory. Through an extensive analysis, we investigate why memory sharing helps. Shared and local memory develop different representations, and later recursions attend mostly to the shared memory, which also acts as a gradient highway to the first recursion.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02383 [cs.LG]
  (or arXiv:2610.02383v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02383

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Giovanni Monea [view email]
[v1] Thu, 1 Oct 2026 19:05:18 UTC (822 KB)

来源:arXiv:cs.LG · arxiv.org