arXiv:cs.LG(机器学习,全量分类)· Xinchen Du, Zhengze Zhou, Wenhui Zhu, Han Yu, Sen Na, Rohit Jain, Alborz Geramifard·· 14 小时前AI 评分33
SHARPO:面向智能体强化学习的段级信用分配机制
SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning
AI 导读
SHARPO 是一种面向智能体强化学习的段级信用分配机制,通过计算每个环境交互段内的教师-学生对数概率差,对 GRPO 优势值施加有界乘数,使信用可在不同段间差异化分配。基于 Qwen2.5-7B-Instruct,SHARPO 在 ALFWorld 和 WebShop 基准上优于 GRPO、SDAR、RLSD 和 StepOPSD 等基线。
正文
Abstract:Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relative Policy Optimization (GRPO) at the level of environment-facing segments. Inspired by the existing on-policy self-distillation (OPSD) method, SHARPO computes teacher-student log-probability gaps within each segment and uses the resulting signal to compute a bounded multiplier on the GRPO advantage. This multiplier is shared by all tokens within the segment, allowing credit to vary across different segments. With Qwen2.5-7B-Instruct, SHARPO outperforms existing baselines on the ALFWorld and WebShop benchmarks, including GRPO, SDAR, RLSD, and StepOPSD.
| Comments: | 13 pages, 3 tables, 2 figures |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.00838 [cs.LG] |
| (or arXiv:2610.00838v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00838 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xinchen Du [view email]
[v1]
Wed, 30 Sep 2026 23:51:11 UTC (1,409 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org