跳到正文
arXiv:cs.LG· Jonas Schweisthal, Yuxin Wang, Athiya Deviyani, Stefan Feuerriegel, Dennis Frauen·· 4 小时前AI 评分34

面向推理时对齐的 Best-of-N 策略高效评估方法

Efficient Best-of-N policy evaluation for inference-time alignment

AI 导读

研究者提出一种仅依赖样本的 Best-of-N(BoN)策略评估与选择框架,无需访问响应似然即可估计密度比。该框架给出双稳健估计器 BoN-DR,可跨候选预算复用共享辅助样本池,并在奖励估计器设定错误时仍保持渐近推断有效。在合成实验与 GSM8K(多种参考模型和奖励模型)上,该方法能准确估计 BoN 策略价值并选出有效采样预算。

正文

View PDF HTML (experimental)

Abstract:Best-of-N (BoN) is a common inference-time alignment method that selects the highest-scoring response among N samples from a reference model. Evaluating BoN policies from logged data is challenging under sample-only access because standard off-policy estimators require density ratios that depend on unavailable response likelihoods. In this paper, we propose a sample-only framework for evaluating and selecting BoN policies without access to these likelihoods. We show that the order-statistic structure of BoN allows the required density ratios to be expressed through score-rank probabilities that are estimable from samples alone. We then develop a doubly robust estimator of the BoN policy value (BoN-DR) that efficiently reuses a shared auxiliary sample pool across candidate budgets. We establish valid asymptotic inference even under reward estimator misspecification and prove the efficiency of our BoN-DR estimator. Since larger budgets can amplify errors in the score function and lead to reward overoptimization, we derive two selection rules: (i) maximizing the estimated policy value and (ii) maximizing a lower confidence bound on the improvement over the reference policy, which accounts for estimation uncertainty and provides a no-harm guarantee. Across synthetic experiments and GSM8K with multiple reference and reward models, our framework accurately estimates BoN policy values and selects effective sampling budgets.
Comments: Preprint
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as: arXiv:2610.09250 [cs.LG]
  (or arXiv:2610.09250v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09250

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Dennis Frauen [view email]
[v1] Wed, 7 Oct 2026 00:26:39 UTC (1,161 KB)

来源:arXiv:cs.LG · arxiv.org