arXiv:cs.LG· Guilhem Loussouarn, Nancy Nayak, Kin K. Leung·· 3 小时前AI 评分37
分阶段强化学习该用单一策略还是多策略?一篇理论分析
Single or Multiple Policies for Phase-Structured Reinforcement Learning?
AI 导读
针对分阶段非平稳强化学习,研究证明共享的单一状态增强策略理论上可达到任意多策略方案的性能,但实际表现取决于函数逼近、学习与优化过程,以及多策略方案的样本效率和策略切换时的连续性损失。
正文
Abstract:Many reinforcement-learning (RL) problems are non-stationary yet structured and can be decomposed into phases, each with its own transition probabilities and reward functions. When the phase sequence is known, the common solution augments the state with information to satisfy the Markovian property and applies standard RL techniques. However, prior work finds that the multi-policy approach for different phases can outperform a single state-augmented policy shared among the phases, for reasons that remain unclear. In this work, we first show that the shared policy can theoretically achieve performance of any multi-policy solution. However, whether a multi-policy solution can perform better than the corresponding single shared policy in practice depends on function approximation, learning and optimization processes, as well as, for multi-policy solutions, the sample efficiency and loss of continuity from one policy to another. We propose a regime-based phase decomposition method to identify which policy can provide better performance. The method is based on consideration of the duration of transient system dynamics relative to the duration of the quasi-stationary period. Numerical experiments are conducted with different non-stationary RL problems to validate our four major hypotheses: (a) longer phase durations favor multi-policies, (b) the heterogeneity between phases increases the burden on single policy, (c) multi-policies need sufficient data for each phase, and (d) environment-specific transition dynamics between phases can affect which policy is preferable.
| Comments: | 40 pages, 13 figures, main paper with appendix |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Systems and Control (eess.SY) |
| Cite as: | arXiv:2610.03475 [cs.LG] |
| (or arXiv:2610.03475v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03475 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Guilhem Loussouarn [view email]
[v1]
Fri, 2 Oct 2026 15:48:01 UTC (446 KB)
来源:arXiv:cs.LG · arxiv.org