跳到正文
原文
elvis· @omarsar0 · X·· 3 小时前AI 评分44
AI 导读

ProVer 提出用 LLM 裁判定位轨迹中决定成败的关键片段,再由 rollout 决定该片段应得多少信用,解决 GRPO 给轨迹中每个 token 相同优势、无法区分关键步骤的问题。在 ALFWorld、WebShop 和 SearchQA 上,该方法相对 GRPO 对 Qwen3.5-2B 提升 9.91%、对 Qwen3.5-4B 提升 7.12%,且裁判换成更小模型时仍有效。

正文

Good paper on credit assignment for agent RL.

The main finding is that you want an LLM judge to choose where to check a trajectory, and the rollouts to decide how much credit that step gets.

GRPO gives every token in a trajectory the same advantage, so the training signal cannot tell the decisive step from the rest.

ProVer has a judge compare successful and failed rollouts and name the segment it thinks caused the difference. It then samples continuations from just before and just after that segment and uses the change in success rate as the segment's advantage.

Across ALFWorld, WebShop and SearchQA, this gives relative improvements over GRPO of 9.91% for Qwen3.5-2B and 7.12% for Qwen3.5-4B. It still helps when the judge is a smaller model.

Paper: https://arxiv.org/abs/2609.36178

Chat with Paper: https://academy.dair.ai/papers/targeting-pivotal-decisions-for-credit-assignment-in-agentic-reinforcement-learn-2609.36178

来源:elvis · x.com