跳到正文
arXiv:cs.LG· Alexey Kurennoy, Ramil Yarullin, Fergal Reid·· 3 小时前AI 评分34

Metropolis-Hastings 在策略组合中优于重要性重采样

Metropolis-Hastings Dominates Importance Resampling for Policy Composition

AI 导读

针对 LLM 解码时策略组合的采样偏差问题,研究证明在任意 rollout 预算下,基于 independence Metropolis-Hastings(MH)的迭代修正所产生的输出分布,按所有凸 f-散度衡量都至少与同等预算的 sampling-importance-resampling(SIR)一样接近目标分布。

正文

View PDF HTML (experimental)

Abstract:Post-training a large language model (LLM) often requires exploring trade-offs between multiple rewards, but retraining for each trade-off is expensive. Decoding-time policy composition allows these trade-offs to be adjusted by combining reward-specific policies at inference time. This composition targets a weighted product of the policies' probabilities over complete responses, but standard implementations combine their next-token probabilities, generally introducing sampling bias. We analyze a known iterative correction based on independence Metropolis-Hastings (MH). Our main result shows that, for every rollout budget, MH produces an output distribution at least as close to the target as sampling-importance-resampling (SIR) with the same budget, as measured by every convex f-divergence. We also derive a lower bound on MH's improvement over the uncorrected decoder in a consensus objective measuring agreement with the supplied policies. We further characterize the correction's sampling error in two asymptotic regimes: when the reward-specific policies approach agreement, and when the log ratio between target and uncorrected-decoder probabilities fluctuates increasingly widely, as can happen for long responses. We complement our analysis with experiments in enumerable and LLM-scale settings.
Comments: 56 pages, 6 figures
Subjects: Machine Learning (cs.LG)
MSC classes: 60J22 (Primary), 65C05, 68T50 (Secondary)
Cite as: arXiv:2610.03480 [cs.LG]
  (or arXiv:2610.03480v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03480

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Alexey Kurennoy [view email]
[v1] Fri, 2 Oct 2026 15:49:38 UTC (330 KB)

来源:arXiv:cs.LG · arxiv.org