跳到正文
原文
OpenBMB· @OpenBMB · X·· 17 小时前AI 评分41
AI 导读

面壁智能联合清华THUNLP提出扩散奖励模型(DRM),不再将人类偏好压缩为单一分数,而是学习完整奖励分布,保留分歧、不确定性与多种合理判断。DRM可利用分布不确定性识别不稳定的奖励决策,分布感知排序在Best-of-N选择上优于取均值;作为RLHF训练奖励时,下游策略表现也超过标量奖励基线。

正文

A reward of “3” can mean two completely different things.
Everyone thinks a response is mediocre — or half the people love it while the other half hate it.
Most Reward Models cannot tell the difference.
Introducing Diffusion Reward Models (DRM): instead of collapsing human preference into a single score, DRM learns the full reward distribution, preserving disagreement, uncertainty, and multiple plausible judgments.
✨ Paper:https://arxiv.org/abs/2609.33803
🤗 Models: https://huggingface.co/Teburile/DRM
💻 GitHub: https://github.com/thunlp/DRM

Why it matters:
Human disagreement is structured, not just noise. On datasets with repeated annotations, judgments often form separated or polarized patterns. More importantly, as human disagreement increases, DRM’s learned reward distribution becomes increasingly multimodal.

The distribution is useful, not just descriptive. DRM can use distributional uncertainty to identify unstable reward decisions, and distribution-aware ranking improves Best-of-N selection beyond simply taking the mean reward.

Reward Models get their own test-time scaling. Instead of only spending more compute on generating more responses, DRM can keep the response fixed and sample its reward distribution more times. More reward samples give a more reliable estimate — a new scaling axis unavailable to deterministic scalar RMs.

And the gains survive RLHF. When used as the training-time reward, DRM improves downstream policy performance over scalar reward baselines, showing that the benefit is not limited to offline RM benchmarks.

The takeaway: Reward modeling may lose something important when it compresses every human judgment into one number.
“Everyone thinks this is average” and “people strongly disagree about this” should not look identical to a Reward Model.
DRM makes that difference visible — and usable.

来源:OpenBMB · x.com