跳到正文
原文
HuggingFace Daily Papers(社区热门论文)·· 13 小时前AI 评分39

BiasReducer:面向奖励模型的自适应偏见缓解框架

BiasReducer: Adaptive Bias Mitigation for Reward Models

AI 导读

BiasReducer 通过只编辑奖励模型的线性奖励头,自动识别并削减模型对长度、自信度等表层属性的依赖,无需重新训练。在五个奖励模型上,BiasReducer-M 在三个基准上平均提升 8.3、18.0 和 6.9 个百分点,优于两个基于训练的基线方法。该改进可迁移至下游,减少冗余啰嗦与谄媚,同时保持相当的评判质量。

正文

Published on Sep 26

Authors:

,

,

,

Abstract

Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods either retrain the reward model or apply a fixed correction to one known bias, such as a preference for longer responses. Retraining requires additional data and computational resources, while existing editing methods require the target bias to be specified in advance and use a fixed edit for that bias. To this end, we propose BiasReducer, a lightweight framework that edits only the linear reward head and selects the relevant edits for each new dataset. First, BiasReducer uses a sparse autoencoder (SAE)-style encoder to learn which attributes (e.g., length and confidence) the reward model is sensitive to. Second, it learns how to reduce the reward model's dependence on each attribute by determining which direction to adjust the reward head and how much to adjust it. Third, for a new dataset, it ranks the attributes by their influence on reward scores, selects the relevant ones, and edits the reward model accordingly. BiasReducer consistently improves reward-model robustness to biases toward superficial response attributes. Across five reward models, BiasReducer-M improves the three benchmarks by 8.3, 18.0, and 6.9 percentage points on average, outperforming the two training-based baselines. The gains transfer downstream, reducing unnecessary verbosity and sycophancy while maintaining comparable judged quality.

View arXiv page View PDF GitHub 1 Add to collection

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.32720 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.32720 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.32720 in a Space README.md to link it from this page.

Collections including this paper 1

来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co