arXiv:cs.LG(机器学习,全量分类)· Ryunyi Lee, Kangjun Noh, Somin Kim, Heedong Kim, Kyungwoo Song·· 14 小时前AI 评分33
Range-GRPO:通过奖励区间成对关系进行策略优化
Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals
AI 导读
研究者提出 Range-GRPO,一种半监督后训练框架,将 LLM-as-a-Judge 的伪奖励表示为经保形校准的奖励区间,并在 GRPO 中直接对奖励区间做成对比较,而非压缩为单点奖励,使区间不确定性同时影响学习信号的方向与幅度。
正文
Abstract:As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-rewards as conformally calibrated reward ranges. We propose Range-GRPO, a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. In Group Relative Policy Optimization (GRPO), learning signals depend on relative reward comparisons within each rollout group. The proposed objective compares reward ranges pairwise rather than reducing them to point rewards, allowing interval uncertainty to affect both the magnitude and direction of these signals. Our theoretical analysis characterizes this distinction and shows that the proposed objective recovers the this http URL advantage when all reward ranges collapse to points. Empirically, Range-GRPO achieves the highest in-distribution and out-of-distribution average performance among the evaluated semi-supervised methods while requiring fewer training resources.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.01548 [cs.LG] |
| (or arXiv:2610.01548v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01548 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ryunyi Lee [view email]
[v1]
Thu, 1 Oct 2026 12:16:34 UTC (322 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org