arXiv:cs.LG(机器学习,全量分类)· Scott W. Viteri (Stanford University), Laura Gomezjurado Gonzalez (Stanford University), Clark Barrett (Stanford University)·· 14 小时前AI 评分37
内在奖励何时才能引导探索?一项基于反事实信息的强化学习探索准则
When Do Intrinsic Rewards Lead to Exploration?
AI 导读
研究提出一种形式化探索准则,通过策略所获取的反事实信息来比较策略优劣,即其历史能在多大程度上替代其他策略下的经验。作者构建了一个简单环境,其中基于计数、预测误差、赋权(empowerment)和信息增益的既有内在奖励目标,其最大化策略在获取反事实信息上均处于帕累托次优。论文还给出这些目标失效的原因、既有内在奖励成功促进最优探索的条件,并构造了一个在探索严格改善时赋予更高价值的新目标。
正文
Abstract:Intrinsic rewards are designed to guide exploration in reinforcement learning by assigning value to an agent's experience, for example through prediction error or learning progress. However, maximizing these rewards need not produce the most informative experience available. We propose a formal criterion for exploration that compares policies by the counterfactual information they acquire: how well their histories can substitute for experience under alternative policies. We construct a single, simple environment in which specified count-based, prediction-error, empowerment, and information-gain objectives have maximizing policies that are Pareto-suboptimal at acquiring counterfactual information. We explain these failures and establish conditions under which existing intrinsic rewards successfully encourage optimal exploration. We also construct an objective that assigns a higher value whenever exploration strictly improves under our criterion.
| Comments: | 45 pages, 4 figures; includes mathematical appendices. Code, data, and Lean proof sources: this https URL |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02159 [cs.LG] |
| (or arXiv:2610.02159v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02159 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Scott Viteri [view email]
[v1]
Thu, 1 Oct 2026 17:52:41 UTC (1,561 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org