跳到正文
arXiv:cs.LG· M\'onika Farsang, Ramin Hasani, Daniela Rus, Radu Grosu·· 2 天前AI 评分37

通过输入重塑与深度循环改进 SSM:LRU、S5、LinOSS、LrcSSM 四种架构一致受益

Reshape and Recur: Improving SSMs with Input Reshaping and Depth Recurrence

AI 导读

研究提出两种正交方法改进状态空间模型(SSM):引入深度循环,使 k 参数循环 SSM 迭代 M 次即可达到 k·L 独立参数标准 SSM 的性能(M ≤ L),进一步降低内存占用;同时通过拼接低维序列元素的时间步、或对高维元素展平并重新分块特征-时间联合维度,以固定时间粒度提升性能。两项扩展在 LRU、S5、LinOSS、LrcSSM 四种 SSM 架构上均取得一致收益。

正文

View PDF HTML (experimental)

Abstract:State Space Models (SSMs) are increasingly deployed in the Edge because they offer, at comparable performance, a smaller memory/training/inference footprint, compared to Large Language Models (LLMs). These three advantages are a direct consequence of the time recurrence inherent in the SSMs architecture. Here, we further improve this recurrent architecture by positively answering two previously underexplored, orthogonal questions: (1) Can we reduce SSMs memory-footprint without any performance penalty, by also employing depth recurrence? (2) Can we increase SSMs performance by using a fixed and consistent time-granularity across all tasks? The first question is somewhat unexpected, given that SSMs are already recurrent. However, the orthogonal depth recurrence further decreases SSMs memory footprint. We show that a looped SSM with $k$ parameters adaptively iterated $M$ times, achieves a performance comparable to a standard SSM with $k \cdot L$ independent parameters, where $M \leq L$. The second question is also unexpected given the time-recurrent nature of the SSMs architecture. However, it makes perfect sense for the time-parallel training of SSMs on the entire input sequence. We show that concatenating time steps for lower-dimensional sequence elements, or flattening and re-chunking the joint feature-time dimension for high-dimensional ones, can improve the baseline by enhancing the way information is presented to the model. Our results for both extensions lead to consistent benefits across four representative SSM architectures: LRU, S5, LinOSS, LrcSSM.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2605.16048 [cs.LG]
  (or arXiv:2605.16048v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2605.16048

arXiv-issued DOI via DataCite

Submission history

From: Mónika Farsang [view email]
[v1] Fri, 15 May 2026 15:18:12 UTC (77 KB)
[v2] Thu, 1 Oct 2026 12:17:44 UTC (271 KB)

来源:arXiv:cs.LG · arxiv.org