arXiv:cs.LG(机器学习,全量分类)· Yujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni Montana·· 7 小时前AI 评分28
最优传输遇上强化学习:一篇综述
Optimal Transport Meets Reinforcement Learning: A Survey
AI 导读
这篇综述系统梳理了最优传输(OT)在强化学习目标与算法中的应用,覆盖模仿学习、离线 RL 和分布偏移部署等分布弱重叠场景。文章对每种方法分析了 OT 所扮演的角色、所比较的分布、采用的 OT 形式及对时间结构的处理,并讨论了代价设计与计算挑战,指出可扩展轨迹级传输、质量不匹配的原则化处理及 OT 正则化 RL 的理论分析等开放问题。
正文
Abstract:Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, commonly used divergences may become ineffective when these distributions overlap weakly, which is frequently encountered in imitation learning, offline RL, and deployment under distribution shift. Optimal transport (OT) offers an alternative by measuring the cost of \emph{moving} probability mass from one distribution to another under a ground cost that encodes task geometry. This survey covers how OT is used inside RL objectives and algorithms. For each method, we identify: the role OT plays, the distributions compared, the OT formulation used, and the treatment of temporal structure. Beyond categorising existing methods, we discuss the motivations behind different OT choices, practical considerations such as cost design and computational challenges, and highlight open problems including scalable trajectory-level transport, principled handling of mass mismatch, and theoretical analysis for OT-regularised RL.
| Subjects: | Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.01413 [stat.ML] |
| (or arXiv:2610.01413v1 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01413 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Giovanni Montana [view email]
[v1]
Thu, 1 Oct 2026 10:19:53 UTC (209 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org