跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Oleg Shchendrigin, Egor Cherepanov, Aleksandr I. Panov, Alexey K. Kovalev·· 15 小时前AI 评分32

ALER:面向强化学习的自适应可学习经验重写

ALER: Adaptive Learnable Experience Rewriting for Reinforcement Learning

AI 导读

针对部分可观测强化学习中记忆需支持重写与经验融合的需求,研究者提出 ALER(Adaptive Learnable Experience Rewriting),将 LSTM 与槽位记忆结合,通过独立寻址的 Gumbel-Softmax 写入覆盖单个槽位,并用学习门控在策略与价值头之前融合检索内容与循环状态。

正文

View PDF HTML (experimental)

Abstract:In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the ability to keep information unchanged until it is needed. We formalize two further requirements. Rewriting sets the decision-relevant content to a value independent of the old one, and experience fusion transforms the old content by a rule that a later observation specifies. For tasks built from such updates, we count the memory states that a solution needs, and several baselines reach their lowest success rates on compositions that need more states. We introduce ALER (Adaptive Learnable Experience Rewriting), an agent that pairs an LSTM with a slot memory. An independently addressed Gumbel-Softmax write that concentrates its weight on one slot overwrites that slot, and a learned gate fuses the retrieved content with the recurrent state before the policy and value heads. We also introduce Rune-Mazes, three environments in which rune observations invert, cancel, reset, or repeat updates of a hidden cue under vector and pixel observations. Against seven baselines, ALER reaches a success rate of at least $0.82$ in all sixteen Endless T-Maze configurations and at least $0.99$ on all five Rune T-Maze compositions, and it has the highest mean success rate on four-branch Rune Multi-Corridor with an Invert rune. On pixel-based Rune MiniGrid Memory, it has a higher mean success rate than PPO-LSTM in eight of ten configurations. Project page: this https URL.
Comments: 28 pages, 12 figures, 18 tables
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.00592 [cs.LG]
  (or arXiv:2610.00592v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00592

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Egor Cherepanov [view email]
[v1] Wed, 30 Sep 2026 18:55:45 UTC (805 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org