arXiv:cs.AI· Xuancheng Li, Beining Wang, Haitao Li, Heng Wang, Yujia Zhou, Qingyi Pan, Blaze Chen, Yiqun Liu, Min Zhang, Qingyao Ai·· 6 小时前AI 评分37
EnGRICH:用人类批判增强生成式奖励模型
EnGRICH: Enhancing Generative Reward Modeling with Critiques from Humans
AI 导读
EnGRICH 是一个生成式奖励模型(GRM)训练框架,通过引入从少量人类批判中学习的 MetaCritic,为生成的批判提供过程奖励和结构化指导。MetaCritic 构建针对具体回复的评分标准,评估批判的证据覆盖度与正确性,并在训练中持续优化以泛化到仅有偏好结果的数。在七个奖励模型基准上,EnGRICH 一致优于竞争基线,推理时训练好的 GRM 可独立运行。
正文
Abstract:Generative reward models (GRMs) are important for LLM optimization. Unlike scalar reward models, GRMs generate natural-language critiques alongside preference judgments, providing finer-grained evaluation signals. Their effectiveness depends heavily on critique reliability. However, existing GRM training typically uses final preference correctness as outcome supervision. Because the preference outcome space is highly constrained, unreliable critiques can still yield correct outcomes and thus be reinforced. Recent work leverages human critiques for process supervision, but such critiques are scarce and are often reduced to scalar rewards, leaving their fine-grained evaluative information underutilized. We argue that evaluative criteria learned from human critiques can be generalized to broader outcome-only preference data. To this end, we propose \textbf{EnGRICH}, a GRM training framework that pairs the GRM with a training-time MetaCritic learned from a small set of human critiques. MetaCritic constructs response-specific rubrics and uses them to evaluate the evidence coverage and correctness of generated critiques. The resulting signals provide both process rewards for fine-grained credit assignment and structured guidance for exploring better critiques. During GRM training, MetaCritic is further optimized to generalize human-grounded evaluative criteria to outcome-only data. At inference, the trained GRM operates independently. Experiments across seven reward-model benchmarks show that EnGRICH consistently improves over competitive baselines, while further analyses validate the effectiveness of its core mechanisms.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.05370 [cs.AI] |
| (or arXiv:2610.05370v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.05370 arXiv-issued DOI via DataCite |
Submission history
From: Xuancheng Li [view email]
[v1]
Sun, 4 Oct 2026 16:51:20 UTC (553 KB)
[v2]
Tue, 6 Oct 2026 16:53:07 UTC (553 KB)
来源:arXiv:cs.AI · arxiv.org