arXiv:cs.LG· Richard Lee Kim, Yeongmin Kim, Gyuwon Sim, Taekyu Kim, Minsang Park, Il-chul Moon·· 6 小时前AI 评分39
BoNG:面向测试时扩散对齐的 Best-of-N 引导方法
Best-of-$N$ Guidance for Test-time Diffusion Alignment
AI 导读
研究者提出 Best-of-N Guidance(BoNG),将 BoN 采样原理直接融入反向扩散过程,通过去噪粒子间的非对称引导交互将粒子群导向高奖励区域。在 36 项对比实验中,BoNG 有 29 项取得最佳表现,对 SMC 与 Vanilla BoN 的排名第一率达 80.56%。该方法还支持多输出,ImageReward 分数为最新基于采样的引导方法的 1.3 倍,并实现 1.6 倍加速。
正文
Abstract:Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-$N$ (BoN) sampling, which draws $N$ i.i.d. samples from a pre-trained diffusion model and outputs the single highest-reward sample. Despite its empirical success, BoN makes limited use of reward information, as it is incorporated only at the final selection stage without influencing the reverse diffusion trajectory during sampling. Consequently, BoN sampling does not improve the average alignment of generated samples and is primarily suited to single-output settings. We propose Best-of-$N$ Guidance (BoNG), a novel method that integrates the principle of BoN sampling directly into the reverse diffusion process. BoNG performs online BoN selection over denoising particles and adjusts the reverse diffusion process to steer the particle population toward higher-reward regions during generation. Specifically, by introducing an asymmetric guidance interaction among denoising particles, BoNG uses the current BoN particle as a guidance signal to the rest of the particle population. This particle-level interaction reshapes the sampling process toward higher-reward regions, enabling BoNG to improve not only the final best sample beyond Vanilla BoN sampling, but also the average quality of generated samples. Over 36 empirical comparisons, BoNG achieves the best performance in 29 cases, ranking first in 80.56% of the comparisons against SMC and Vanilla BoN sampling. BoNG also supports multi-output capability, achieving 1.3$\times$ ImageReward score of the latest sample-based guidance method with a 1.6$\times$ speedup. We release the code at this https URL.
| Comments: | NeurIPS 2026 |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.05108 [cs.LG] |
| (or arXiv:2610.05108v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.05108 arXiv-issued DOI via DataCite |
Submission history
From: Richard Lee Kim [view email]
[v1]
Sun, 4 Oct 2026 10:35:32 UTC (40,380 KB)
[v2]
Tue, 6 Oct 2026 07:19:42 UTC (2,042 KB)
来源:arXiv:cs.LG · arxiv.org