arXiv:cs.AI· Arip Asadulaev, Aladin Djuhera, Karim Salta, Holger Boche, Fakhri Karray, Martin Takac·· 5 小时前AI 评分45
Tropical Reinforcement Learning:用热带半环改造强化学习提升 LLM 组合推理
Tropical Reinforcement Learning
AI 导读
研究者提出 Tropical Reinforcement Learning,将强化学习中替代方案概率的“求和”改为“取最大值”,引入热带半环,使状态价值对应最可能被验证解的对数概率及可复用路径。配套训练算法 TROPIC 面向确定性、可重置且结果可验证的环境,在 Sokoban、Countdown、FrozenLake、WebShop 四个智能体任务上比最强同策略基线高出最多 16 个百分点。
正文
Abstract:Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities sum to one, reinforcing one solution can make the model forget another that was never shown to be wrong. This makes expected return a poor fit for compositional reasoning, where a solution must be assembled from reasoning steps that the model produces in separate, often failed, attempts but rarely produces together. To address this, we propose Tropical Reinforcement Learning, which rests on a simple change of algebra: instead of adding the probabilities of alternative solutions, we take their maximum, which yields the tropical semiring. The value of a state then becomes the log-probability of its most likely verified solution, together with an explicit path that can be replayed and reused. This enables true composition, since the best prefix and the best suffix meeting at a shared state can be joined even when they come from different rollouts. To put this into practice, we introduce TROPIC, a training algorithm for deterministic, resettable environments with verifiable outcomes. On four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points. Changing the algebra of reinforcement learning, not just its estimators, can thus substantially improve compositional reasoning in language models
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02478 [cs.AI] |
| (or arXiv:2610.02478v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02478 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Arip Asadulaev [view email]
[v1]
Thu, 1 Oct 2026 20:56:11 UTC (1,464 KB)
来源:arXiv:cs.AI · arxiv.org