arXiv:cs.LG· Amir Esterhuysen, Anders Jonsson·· 7 小时前AI 评分38
强化学习中的终端表示(Terminal Representation)
The Terminal Representation in Reinforcement Learning
AI 导读
论文提出终端表示(TR),一种结构上不同于后继表示(SR)和默认表示(DR)的强化学习表示学习方法,可学习为更低维对象并直接用于选项发现、奖励塑形、迁移学习和探索,无需特征向量计算,且能绕过特征分解对对称转移动态的假设。TR 被证明嵌入于 DR 的顶部特征向量中,可在不做特征分解的情况下捕获相同知识,并以更低的计算开销进行学习、存储和使用。
正文
Abstract:Representation learning is a powerful tool for spatio-temporal abstraction within reinforcement learning (RL). Two well established approaches are through the successor representation (SR) and the default representation (DR). The SR encodes states by the future trajectories they induce, capturing information flow decoupled from reward. The DR builds on this by weighting trajectories with reward, integrating credit-assignment structure into the representation. Eigenvectors of both representations have been used to support a range of downstream tasks -- including option discovery, reward shaping, transfer learning, and exploration. We introduce a structurally distinct formulation: the terminal representation (TR). The TR encodes reward-weighted trajectories similarly to the DR, but can be learned as a lower-dimensionality object, and can be used directly for the mentioned applications without eigenvector computations. Eigendecomposition also imposes the assumption of symmetric transition dynamics, which the TR can bypass. In this work we develop the theoretical foundations of the TR: its derivation, convergence of two learning algorithms, its use for zero-shot compositionality, and equivalences between alternative reward formulations. We further show the TR is embedded in the top DR eigenvector, allowing it to capture the same underlying knowledge without eigendecomposition. Additionally, we provide empirical evidence of the TR as a viable alternative to existing representations in subsidiary applications, while requiring less computational overhead to learn, store, and use.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2605.31289 [cs.LG] |
| (or arXiv:2605.31289v3 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.31289 arXiv-issued DOI via DataCite |
Submission history
From: Amir Esterhuysen [view email]
[v1]
Fri, 29 May 2026 13:24:28 UTC (1,068 KB)
[v2]
Fri, 17 Jul 2026 15:05:43 UTC (1,060 KB)
[v3]
Tue, 6 Oct 2026 14:15:29 UTC (1,090 KB)
来源:arXiv:cs.LG · arxiv.org