arXiv:cs.AI· Hongqiang Lin, Dongxu Zhang, Yiding Sun, Mingzhe Li, Ning Yang, Haijun Zhang·· 4 小时前AI 评分26
PSPO:基于后验采样的离线策略优化
Offline Policy Optimization with Posterior Sampling
AI 导读
研究者提出 PSPO,将动力学模型视为随机变量而非点估计,通过交替更新后验分布与策略,实现对 OOD 区域的可控探索,并设计了带收敛保证的正则化优化算法。在标准 benchmark 上,PSPO 性能优于当前最优基线,且兼具无过度悲观特性与鲁棒性。
正文
Abstract:A fundamental challenge in model-based offline reinforcement learning (RL) lies in the trade-off between generalization and robustness against exploitation errors in out-of-distribution (OOD) regions. The key to resolving this trade-off lies in enabling the model to explore OOD regions that remain consistent with underlying physical dynamics. However, achieving this is challenging because limited data cannot uniquely identify the dynamics model, and unconstrained exploration is risky. Existing methods often overlook this nuance, addressing the risk through excessive pessimistic regularization, which ensures robustness but sacrifices generalization. To address this, we propose PSPO, which treats the dynamics model as a random variable rather than a point estimate. This formulation inherently allows for controlled exploration of OOD regions. By alternately updating the posterior distribution and the policy, we design a regularized optimization algorithm with convergence guarantees. Experiments on standard benchmarks demonstrate that PSPO achieves superior performance compared to state-of-the-art baselines. Further analysis confirms that our method attains the desired pessimism-free property while maintaining robustness, and ablation studies verify the effectiveness of each proposed module.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2605.07393 [cs.AI] |
| (or arXiv:2605.07393v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2605.07393 arXiv-issued DOI via DataCite |
Submission history
From: Hongqiang Lin [view email]
[v1]
Fri, 8 May 2026 07:48:21 UTC (2,517 KB)
[v2]
Fri, 2 Oct 2026 07:16:08 UTC (2,537 KB)
来源:arXiv:cs.AI · arxiv.org