跳到正文
arXiv:cs.LG· Syed Ibrahim Omer, Ginny Y. Wong. Xiangyu Zhao·· 4 小时前AI 评分35

MaRK:面向状态空间模型动态算子条件化的马尔可夫自适应循环核

MaRK: Markov-adapted Recurrent Kernels for Dynamic Operator Conditioning in State Space Models

AI 导读

MaRK 是一种动态算子条件化框架,可将上下文向量直接映射为冻结 SSM 的循环(A)、读入(B)、读出(C)、跳跃(D)和离散化(Δ)参数的有界调制。在冻结的 111M 参数 Hydra SSM 上,Hypernet、Chebyshev 多项式与 DCT 三种适配器仅需 6.3–11M 可训练辅助参数即可从双向目标过渡到迭代扩散机制,其中 Chebyshev 变体验证损失最低,为 2.55。

正文

View PDF HTML (experimental)

Abstract:State Space Models (SSMs) offer an efficient alternative to Transformers for sequence modeling, yet conditioning pre-trained SSMs for iterative generation typically operates outside the recurrent operator, through input injection or activation modulation. While such mechanisms expose the model to conditioning information, they leave the underlying temporal dynamics fixed. We introduce MaRK (Markov-adapted Recurrent Kernels), a dynamic operator-conditioning framework that maps context vectors directly into bounded modulations of a frozen SSM's recurrence ($A$), read-in ($B$), read-out ($C$), skip ($D$), and discretization ($\Delta$) parameters. Viewed through the lens of LPV-SSM systems, MaRK induces a context-indexed family of Markov parameter sequences, allowing each diffusion timestep to reshape the model's input-output memory kernel. We instantiate MaRK on a frozen 111M-parameter Hydra SSM backbone and study three adapter geometries: Hypernet, Chebyshev polynomial, and Discrete Cosine Transform kernels. Since these adapters modify the Markov parameter sequence through low-rank auxiliary maps on the frozen backbone, parameter-efficient fine-tuning arises as a structural consequence of the adaptation mechanism itself, requiring only 6.3--11M trainable auxiliary parameters to transition from a bidirectional objective to an iterative diffusion regime. The bounded recurrence parameterization further yields an analytic Affine Quadratic Stability certificate for the modulated recurrence. Through synthetic LPV recovery experiments and Markov-operator diagnostics, we show that MaRK recovers coordinate-invariant temporal operators under matched assumptions and produces distinct, stable timestep-conditioned memory profiles. Empirically, the Chebyshev variant yields the strongest performance, achieving an average validation loss of 2.55, followed by the DCT (2.59) and Hypernet (3.77) geometries.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.09092 [cs.LG]
  (or arXiv:2610.09092v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09092

arXiv-issued DOI via DataCite (pending registration)

Journal reference: NeurIPS 2026

Submission history

From: Syed Ibrahim Omer [view email]
[v1] Tue, 6 Oct 2026 20:46:26 UTC (1,708 KB)

来源:arXiv:cs.LG · arxiv.org