跳到正文
arXiv:cs.LG· Kaicheng Xiao, Haotian Li, Liran Dong, Guoliang Xing·· 7 小时前AI 评分46

RAM-Net:用稀疏可寻址状态实现线性时间序列建模

RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State

AI 导读

RAM-Net 将循环状态组织为固定大小的独立槽位数组,并用 Address Decoder 把每个 key 或 query 映射为稀疏地址,每步只读写少量槽位以抑制 token 间干扰,且单步状态访问量只取决于所选槽位数而非状态总大小。

正文

View PDF HTML (experimental)

Abstract:Linear attention offers an efficient alternative to full attention with a fixed-size recurrent state. However, this state is shared by all tokens, so information from distinct tokens becomes superposed within it and produces inter-token interference that degrades long-range fine-grained recall. To address this issue, we propose RAM-Net, which replaces dense access to a shared state with sparse address-based access. RAM-Net organizes the recurrent state as a fixed-size array of independent slots and uses an Address Decoder that maps each key or query into a sparse address, selecting a small subset of slots to write to or read from at each step. This design directs tokens with non-overlapping addresses to disjoint slots, suppressing inter-token interference, while keeping per-step state access dependent only on the number of selected slots rather than the total state size. Empirically, RAM-Net outperforms strong recurrent baselines on fine-grained long-range retrieval and achieves the lowest perplexity with competitive commonsense reasoning. It does so while accessing fewer state elements per step than all baselines, e.g., $8\times$ fewer than Mamba2.
Comments: Accepted at NeurIPS 2026. Project page: this https URL
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as: arXiv:2602.11958 [cs.LG]
  (or arXiv:2602.11958v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2602.11958

arXiv-issued DOI via DataCite

Submission history

From: Kaicheng Xiao [view email]
[v1] Thu, 12 Feb 2026 13:55:29 UTC (2,009 KB)
[v2] Tue, 6 Oct 2026 16:19:58 UTC (3,018 KB)

来源:arXiv:cs.LG · arxiv.org