arXiv:cs.LG(机器学习,全量分类)· Sathya Kamesh Bhethanabhotla, Efstratios Gavves, Andr\'e Biedenkapp·· 5 小时前AI 评分36
用复杂值记忆做强化学习:uRNN 作为 PPO 的即插即用替代
Reinforcement Learning with Complex (valued) Memories
AI 导读
研究者将 Unitary 循环网络(uRNN)作为循环 PPO 架构的即插即用替代,提出三种版本,利用复数隐状态的相位自由度提升长期记忆能力。在 rocksample 和 Craftax 等任务上,该方法奖励达到基线的 2-3 倍,并探索了类似量子态测量的相位感知策略。代码已开源,论文被第 19 届欧洲强化学习研讨会 2026 接收。
正文
Abstract:Partially observable environments pose a fundamental challenge in deep reinforcement learning, requiring agents to compress temporal information from observations and maintain a memory to make effective decisions. While there exist many approaches ranging from gated recurrence to attention mechanisms and model-based RL, the search for effective representational techniques that can capture long-term dependencies remains an active area of research. In this work we revisit Unitary recurrent networks (uRNNs) [Arjovsky et al., 2016, Jing et al., 2017], that demonstrated superior gradient flow and associative recall, expressing the recurrence and the hidden state in a complex vector space. Their norm preserving unitary dynamics enable information propagation through long sequences. To this end, we propose three different versions of uRNNs as drop-in replacements for recurrent PPO architectures, and demonstrate that the simple recurrence and the added degree of freedom from the phase of the complex representations enable significant gains over baselines on several memory-improvable tasks, including continuous control. We further explore how to preserve the phase information of the complex hidden state for a phase-aware policy by drawing a parallel to how quantum states are measured. With our methods reaching up to 2-3 $\times$ the reward in environments like rocksample and Craftax compared to the baselines, this work points towards an exciting new direction of representations for RL and the problem of partial observability. Code is available at: this https URL
| Comments: | 19 pages, 4 figures, 9 tables Accepted at the 19th European Workshop on Reinforcement Learning 2026 |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.38598 [cs.LG] |
| (or arXiv:2609.38598v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38598 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sathya Kamesh Bhethanabhotla [view email]
[v1]
Tue, 29 Sep 2026 21:59:34 UTC (2,262 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org