arXiv:cs.LG· Yining Li, Dongchen Han, Jie Fu, Gao Huang·· 4 小时前AI 评分32
DeltaTTT:面向非线性循环记忆的逐层优化
DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory
AI 导读
研究者提出 DeltaTTT,将两层记忆网络的联合内循环优化改为逐层学习,每层分配局部预测目标并用状态相关的 delta 规则更新。该方法在保留非线性读出的同时支持分块并行计算,在 DeltaNet 和 LaCT 骨干上提升了语言建模与检索表现。
正文
Abstract:Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state. Intuitively, this state dependence should allow each update to account for what the memory has already learned and better incorporate new information. However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outperforms its serial counterpart. Our exploratory experiments point to a key underlying difficulty: nonlinear memories can be harder to optimize than linear ones within a single pass over the sequence. To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop optimization of a two-layer memory network with layerwise learning. Each layer is assigned a local prediction target and updated through a state-dependent delta rule. This formulation retains a nonlinear readout while enabling chunkwise parallel computation. Experiments on DeltaNet and LaCT backbones show improvements in language modeling and retrieval over their recurrent baselines.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.08553 [cs.LG] |
| (or arXiv:2610.08553v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08553 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yining Li [view email]
[v1]
Tue, 6 Oct 2026 15:40:06 UTC (393 KB)
来源:arXiv:cs.LG · arxiv.org