跳到正文
arXiv:cs.CL· Xin Guan, Xiaomeng Hu, Shen Huang, Zhenyi Wang, Bo Zhang, Zijian Li, Pengjun Xie, Bo Liu, Jiuxin Cao·· 3 小时前

EvoRubric:自演进评分标准驱动的开放生成强化学习框架

EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation

AI 导读

EvoRubric 是一个协同演进强化学习框架,通过共享策略同时充当推理器与评分标准生成器,结合冻结的 Meta-Verifier 和 Grader 完成开放式生成的评分标准学习。在 Medical、Writing、Science 三类共五个 benchmark 上,8B 模型均分 56.28、14B 模型均分 61.13,较最强基线分别高出 3.09 和 2.06 分。

正文

View PDF HTML (experimental)

Abstract:Reinforcement Learning (RL) has advanced Large Language Models (LLMs) in verifiable domains, while open-ended generation remains challenging due to the absence of definitive rewards. Rubric-based RL provides explicit evaluation criteria, but learning to construct these criteria remains challenging when final-answer correctness is not verifiable. We propose EvoRubric, a co-evolutionary RL framework that combines criterion-validity feedback, response discrimination, and peer agreement to learn rubrics for open-ended generation. A shared policy acts as both a Reasoner and a Rubric Generator, using its current responses and historical rubrics to discover new evaluation dimensions. To combine adaptive rubric discovery with a stable validity check, a frozen copy of the initial policy serves as the Meta-Verifier, while a frozen Grader scores responses against the retained criteria. Discriminative feedback, Leave-One-Out peer consensus, and a persistent memory pool transform this feedback into complementary rewards that jointly optimize both policy roles, closing the loop between response improvement and rubric discovery. EvoRubric improves over matched static and external evolving-rubric baselines across five benchmarks in Medical, Writing, and Science. Across three training seeds, it achieves five-benchmark averages of 56.28 at 8B and 61.13 at 14B, exceeding the strongest matched baselines by 3.09 and 2.06 points, respectively. Human audits assess criterion validity and response quality, and experiments with expert-initialized rubrics demonstrate compatibility with human priors.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2605.29847 [cs.CL]
  (or arXiv:2605.29847v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2605.29847

arXiv-issued DOI via DataCite

Submission history

From: Xin Guan [view email]
[v1] Thu, 28 May 2026 12:28:49 UTC (697 KB)
[v2] Thu, 8 Oct 2026 03:01:22 UTC (538 KB)

来源:arXiv:cs.CL · arxiv.org