arXiv:cs.LG· Matthew Brun, Xu Andy Sun·· 3 小时前AI 评分33
成功条件化策略优化的收敛性研究
On the Convergence of Success Conditioning for Policy Optimization
AI 导读
研究证明成功条件化在广泛类别的马尔可夫决策过程(MDP)中收敛至最优策略,并推导了常见设定下的收敛速率。在折扣 MDP 中,达到 ε-最优策略需 O(1/ε^p) 次迭代,指数 p 取决于问题数据;在单周期 MDP 中,仅需 O(log(1/ε)) 次迭代。
正文
Abstract:Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common to many reinforcement learning applications, yet its limiting behavior and convergence rates are not well understood. In this work, we demonstrate that success conditioning converges to an optimal policy on a broad class of Markov decision processes (MDPs). We also derive convergence rates in some common settings. For discounted MDPs, we prove convergence within $\mathcal{O}(1/\varepsilon^p)$ iterations to an $\varepsilon$-optimal policy, where the exponent $p$ depends on problem data. For single-period MDPs, such a policy is obtained within $\mathcal{O}(\log(1/\varepsilon))$ iterations.
| Subjects: | Machine Learning (cs.LG); Optimization and Control (math.OC) |
| MSC classes: | 90C40 |
| Cite as: | arXiv:2610.03642 [cs.LG] |
| (or arXiv:2610.03642v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03642 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Matthew Brun [view email]
[v1]
Fri, 2 Oct 2026 17:29:38 UTC (19 KB)
来源:arXiv:cs.LG · arxiv.org