跳到正文
arXiv:cs.LG· Soumya Nasipuri, Sayak Ray Chowdhury, Sanjukta Roy·· 4 小时前AI 评分34

对齐中线性奖励的公理可满足性

Axiom Satisfiability of Linear Rewards in Alignment

AI 导读

研究提出放宽线性奖励模型、允许每个候选存在松弛量,以最小总松弛满足带 margin η 的 PO 和 PMC 公理,且不对投票者或比较数据收集方式作假设。当候选数为 m、η 不超过 O(1/m²) 时,最优总松弛上界为 O(1),并给出紧致性实例。新方法对每次错误比较惩罚线性部分,同时最小化总松弛与违规数,实验显示其线性奖励优于线性 BTL。

正文

View PDF HTML (experimental)

Abstract:Learning from human preference data is the dominant route to aligning language models with human values. In linear social choice, where rewards are linear in a fixed feature representation of prompt-response pairs, Ge et al.[2024] show that fitting such a reward by minimizing any non-decreasing convex loss, including BTL, fails PO and PMC. Moreover, no rule that reads only the majority relation can satisfy PO once the output is required to be linearly induced. We ask what it costs to enforce these axioms anyway. To this end, we relax the linear model to allow per-candidate slack. We compute the relaxed linear reward with the smallest total slack that satisfies the axioms with a margin $\eta$, the minimum required difference between two reward values. Our solution satisfies the axioms under no assumptions about the voters or how comparisons were collected. We bound the optimal total slack by $O(1)$ when $\eta$ is at most $O(\frac{1}{m^2})$ for $m$ candidates. Furthermore, we exhibit an instance that forces this bound, concluding that the rate is tight up to constants. In practice, the no. of candidates far exceeds the feature dimension, and only a linear reward can be evaluated on unseen responses. We therefore introduce a new method that charges the linear part for each comparison it gets wrong. It simultaneously minimizes the total slack and the no. of violations, with a parameter $\lambda$ trading off between them. We show that the total slack is monotone but saturating in $\lambda$: raising it reduces the violations of the deployed linear reward and increases the slack, yet the slack stays below $(\eta+\Delta\sqrt{d})\lfloor m^2/4\rfloor$, where $\Delta$ and $d$ are the diameter and dimension of the features, respectively. Experiments on both synthetic and real-life preference data corroborate our theory and show that the linear reward output by our method beats linear BTL.
Comments: 22 pages, 5 figures
Subjects: Computer Science and Game Theory (cs.GT); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
Cite as: arXiv:2610.06892 [cs.GT]
  (or arXiv:2610.06892v1 [cs.GT] for this version)
  https://doi.org/10.48550/arXiv.2610.06892

arXiv-issued DOI via DataCite

Submission history

From: Soumya Nasipuri Mr. [view email]
[v1] Sat, 26 Sep 2026 15:24:29 UTC (110 KB)

来源:arXiv:cs.LG · arxiv.org