跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Jude Waide, Robert Lieck·· 7 小时前AI 评分33

Decision Titan:用测试时训练为离线强化学习带来长期记忆

Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning

AI 导读

研究者将测试时训练(TTT)层引入 Decision Transformer,提出 Decision Titan,首次把 TTT 框架用于离线强化学习。在 X-Maze 环境中,该模型能学习比上下文窗口长 20 倍的长期依赖,并可泛化到训练数据 1.7 倍的长度。但时间泛化依赖所用的时间嵌入,长期依赖的学习能力取决于相关信息如何编码。

正文

View PDF HTML (experimental)

Abstract:Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has proposed tackling this problem with the Test-Time Training (TTT) framework, which stores episodic memories in the parameters of a neural network through gradient descent at both train and test-time. This approach has seen success in the domain of Natural Language Processing, however, to the best of our knowledge it has not yet been applied to the domain of Reinforcement Learning (RL), nor has there been a study analysing how this memory practically functions. In this paper, we study the potential of the TTT framework for offline RL by augmenting a Decision Transformer with TTT layers, dubbed the Decision Titan. We analyse performance and properties of the model in the X-Maze environment, an extension of T-Maze designed to test sequential memory, and investigate how the memory mechanism learns by visualising gate values over time. Our key findings are that Decision Titan can learn long-term dependencies with ranges 20x longer than the context window, generalises to lengths 1.7x the training data, but crucially temporal generalisation depends on the time embeddings used, and the ability to learn long-term dependencies depends on how the relevant information is encoded.
Comments: Accepted at ICML 2026 Workshop on Decision-Making from Offline Datasets to Online Adaptation: Black-Box Optimization to Reinforcement Learning
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.01513 [cs.AI]
  (or arXiv:2610.01513v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.01513

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jude Waide Mr [view email]
[v1] Thu, 1 Oct 2026 11:48:32 UTC (3,412 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org