跳到正文
arXiv:cs.LG· Seyedarmin Azizi, Erfan Baghaei Potraghloo, Minoo Ahmadi, Souvik Kundu, Massoud Pedram·· 5 小时前AI 评分46

Power-SMC:面向免训练 LLM 推理的低延迟序列级幂采样

Power-SMC: Low-Latency Sequence-Level Power Sampling for Training-Free LLM Reasoning

AI 导读

Power-SMC 是一种免训练采样方法,直接瞄准序列级幂分布,推理延迟接近标准解码,在 MATH500、GSM8K、GPQA 和 HumanEval 上准确率持平或超过 Metropolis-Hastings 采样,推理速度最高提升 17.6 倍。

正文

View PDF HTML (experimental)

Abstract:Reasoning ability in large language models is often attributed to \emph{distribution sharpening}: concentrating output probability on high-likelihood sequences. Recent works show that this sharpening effect can be obtained at inference time, without modifying model parameters, and can elicit strong reasoning performance. A natural formalization is the \emph{sequence-level power distribution}, which is proportional to the model's probability raised to an exponent $\alpha>1$. Prior work leveraged Metropolis--Hastings (MH) sampling to draw samples from this distribution and achieves strong results, however, at order-of-magnitude inference slowdowns. We introduce \textbf{Power-SMC}, a \textit{`training-free'} sampling method that targets the same power distribution yielding close to standard decoding latency. Power-SMC maintains multiple candidate sequences in parallel. Each candidate sequence is assigned a score, namely the \emph{importance weight}, that measures how well it matches the power distribution. It then periodically prunes low-scoring candidate sequences in favor of high-scoring ones. We further provide a theoretical justification for the design choices in Power-SMC. Among all next-token sampling strategies that do not rely on future tokens, we prove that sampling temperature $\tau{=}1/\alpha$ uniquely eliminates per-step weight variance. Finally, we characterize the remaining source of weight instability and introduce a gradual sharpening schedule to reduce weight collapse, while targeting the same power distribution. Extensive evaluations on MATH500, GSM8K, GPQA, and HumanEval show that Power-SMC matches or exceeds MH sampling in accuracy, preserves output diversity unlike RL-finetuned models, while \textbf{accelerating inference speed by up to} {$\mathbf{17.6}\times$}. The code is available at this https URL.
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG)
Cite as: arXiv:2602.10273 [stat.ML]
  (or arXiv:2602.10273v3 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2602.10273

arXiv-issued DOI via DataCite

Submission history

From: Seyedarmin Azizi [view email]
[v1] Tue, 10 Feb 2026 20:31:40 UTC (45 KB)
[v2] Mon, 23 Mar 2026 17:01:31 UTC (45 KB)
[v3] Fri, 2 Oct 2026 07:03:09 UTC (972 KB)

来源:arXiv:cs.LG · arxiv.org