跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Xinchen Du, Zhengze Zhou, Wenhui Zhu, Han Yu, Sen Na, Rohit Jain, Alborz Geramifard·· 14 小时前AI 评分33

SHARPO:面向智能体强化学习的段级信用分配机制

SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning

AI 导读

SHARPO 是一种面向智能体强化学习的段级信用分配机制,通过计算每个环境交互段内的教师-学生对数概率差,对 GRPO 优势值施加有界乘数,使信用可在不同段间差异化分配。基于 Qwen2.5-7B-Instruct,SHARPO 在 ALFWorld 和 WebShop 基准上优于 GRPO、SDAR、RLSD 和 StepOPSD 等基线。

正文

View PDF HTML (experimental)

Abstract:Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitation, we introduce Segment-level Hindsight Advantage Reweighting for Policy Optimization (SHARPO), a credit-assignment mechanism that refines Group Relative Policy Optimization (GRPO) at the level of environment-facing segments. Inspired by the existing on-policy self-distillation (OPSD) method, SHARPO computes teacher-student log-probability gaps within each segment and uses the resulting signal to compute a bounded multiplier on the GRPO advantage. This multiplier is shared by all tokens within the segment, allowing credit to vary across different segments. With Qwen2.5-7B-Instruct, SHARPO outperforms existing baselines on the ALFWorld and WebShop benchmarks, including GRPO, SDAR, RLSD, and StepOPSD.
Comments: 13 pages, 3 tables, 2 figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.00838 [cs.LG]
  (or arXiv:2610.00838v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00838

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xinchen Du [view email]
[v1] Wed, 30 Sep 2026 23:51:11 UTC (1,409 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org