arXiv:cs.LG· Alexey Kurennoy, Ramil Yarullin, Fergal Reid·· 3 小时前AI 评分34
Metropolis-Hastings 在策略组合中优于重要性重采样
Metropolis-Hastings Dominates Importance Resampling for Policy Composition
AI 导读
针对 LLM 解码时策略组合的采样偏差问题,研究证明在任意 rollout 预算下,基于 independence Metropolis-Hastings(MH)的迭代修正所产生的输出分布,按所有凸 f-散度衡量都至少与同等预算的 sampling-importance-resampling(SIR)一样接近目标分布。
正文
Abstract:Post-training a large language model (LLM) often requires exploring trade-offs between multiple rewards, but retraining for each trade-off is expensive. Decoding-time policy composition allows these trade-offs to be adjusted by combining reward-specific policies at inference time. This composition targets a weighted product of the policies' probabilities over complete responses, but standard implementations combine their next-token probabilities, generally introducing sampling bias. We analyze a known iterative correction based on independence Metropolis-Hastings (MH). Our main result shows that, for every rollout budget, MH produces an output distribution at least as close to the target as sampling-importance-resampling (SIR) with the same budget, as measured by every convex f-divergence. We also derive a lower bound on MH's improvement over the uncorrected decoder in a consensus objective measuring agreement with the supplied policies. We further characterize the correction's sampling error in two asymptotic regimes: when the reward-specific policies approach agreement, and when the log ratio between target and uncorrected-decoder probabilities fluctuates increasingly widely, as can happen for long responses. We complement our analysis with experiments in enumerable and LLM-scale settings.
| Comments: | 56 pages, 6 figures |
| Subjects: | Machine Learning (cs.LG) |
| MSC classes: | 60J22 (Primary), 65C05, 68T50 (Secondary) |
| Cite as: | arXiv:2610.03480 [cs.LG] |
| (or arXiv:2610.03480v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03480 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Alexey Kurennoy [view email]
[v1]
Fri, 2 Oct 2026 15:49:38 UTC (330 KB)
来源:arXiv:cs.LG · arxiv.org