arXiv:cs.LG· Sayak Chakrabarti, Sathish Reddy Indurthi·· 4 小时前AI 评分33
RELACE:面向长程语言智能体的回顾式似然动作信用估计
RELACE: retrospective likelihood-based action credit estimation for long-horizon language agents
AI 导读
RELACE 是一个无需 critic 的框架,将回顾式动作评估与状态条件优势估计结合,用于长程语言智能体的信用分配。它通过 teacher-forced 似然评分比较动作在原始上下文与结果增强上下文下的差异,生成轨迹归一化的回顾因子来重加权折扣回报,无需辅助价值模型、奖励模型或额外自回归 rollout。
正文
Abstract:Group Relative Policy Optimization (GRPO) avoids a separate critic by estimating advantages from rollout groups. For multi-turn agents, however, trajectory-level supervision provides coarse, noisy credit: terminal rewards do not locate errors and can penalize useful actions alongside mistakes. Group-in-Group Policy Optimization (GiGPO) and subsequent methods refine supervision through state-conditioned comparisons, but their credit estimates remain sensitive to downstream decisions and outcomes. We introduce RELACE, Retrospective Likelihood-based Action, a critic-free framework that integrates retrospective action assessment with state-conditioned advantage estimation. RELACE evaluates executed actions through teacher-forced likelihood scoring under both their original contexts and outcome-augmented contexts. Comparing these likelihoods yields a trajectory-normalized retrospective factor that captures outcome-dependent changes in action plausibility, rather than hindsight plausibility alone. We use this factor to reweight discounted task returns and construct local advantages by comparing weighted returns among actions from equivalent states within a task. This couples retrospective relevance with observed reward, producing fine-grained credit that complements trajectory-level GRPO supervision. Temporal smoothing and success-protecting masking further stabilize the local signal. RELACE requires neither auxiliary value nor reward models nor additional autoregressive rollouts for credit estimation. Experiments on ALFWorld and WebShop with Qwen2.5-1.5B-Instruct and Qwen2.5-7B-Instruct demonstrate substantial improvements over GRPO, GiGPO, and HCAPO. With the 1.5B model, RELACE achieves $96.35\%$ success on ALFWorld and $79.43\%$ on WebShop, surpassing GiGPO by $5.47$ and $5.60$ percentage points, respectively.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07349 [cs.LG] |
| (or arXiv:2610.07349v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07349 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sayak Chakrabarti [view email]
[v1]
Mon, 5 Oct 2026 20:19:04 UTC (126 KB)
来源:arXiv:cs.LG · arxiv.org