跳到正文
arXiv:cs.CL· Mingqing Yuan (Soochow University), Xiaobo Liang (Soochow University), Junwei Yang (University of Cambridge), Ziwei Chen (Chalmers University of Technology), Zeren Zhang (Peking University), Hejin Wang (Tsinghua University), Yubin Wang (The Hong Kong University of Science,Technology), Juntao Li (Soochow University)·· 3 小时前AI 评分34

LatentGRM:通过语义保持压缩实现高效生成式奖励建模

Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression

AI 导读

LatentGRM 是一种基于语义分块、压缩与重建的隐空间评估框架,利用评分标准引导压缩,学习支持自主成对判断的紧凑连续轨迹,无需生成文本评估。在相同训练数据和骨干下,LatentGRM 在 4B 和 8B 规模上的偏好准确率与显式 SFT 评判模型相当;LatentGRM-8B 在四个基准领域将评估轨迹压缩 8.9–9.2 倍,vote@5 下评判推理总时间减少 6.1–7.0 倍。

正文

View PDF HTML (experimental)

Abstract:Reward modeling often requires jointly representing and reasoning over multiple evaluation criteria, yet verbalizing this process token by token can incur substantial inference cost. Recent work on latent reasoning suggests that continuous states may support this computation more compactly. We introduce LatentGRM, a latent evaluation framework built on semantic chunking, compression, and reconstruction. By using the structure of rubric-guided evaluations to guide compression, LatentGRM learns compact continuous trajectories that support autonomous pairwise judgments without generating textual assessments. A separate interpreter reconstructs evaluation text from these trajectories, providing an offline view of the information retained under compression. Under matched training data and backbones, LatentGRM achieves competitive aggregate preference accuracy relative to explicit Supervised Fine-Tuning (SFT) judges at both 4B and 8B scales. Across four benchmark domains, LatentGRM-8B compresses evaluation trajectories by 8.9--9.2x and reduces total judge inference time by 6.1--7.0x at vote@5. Controlled rubric interventions show that criterion-dependent preference information is carried through the latent sequence. Together, these results demonstrate that continuous latent evaluation can substantially reduce inference cost while preserving competitive judgment quality.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.09788 [cs.CL]
  (or arXiv:2610.09788v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.09788

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mingqing Yuan [view email]
[v1] Wed, 7 Oct 2026 10:05:06 UTC (326 KB)

来源:arXiv:cs.CL · arxiv.org