跳到正文
arXiv:cs.AI· Phillip Howard, Xin Su, Allen Roush, Manikandan Ravikiran, Runyan Tan, Amir Abdullah·· 6 小时前AI 评分38

拒绝门控解码:高温采样下保持 LLM 拒绝行为

Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling

AI 导读

研究者提出拒绝门控解码(RGD),一种顺序解码方法,可在高温采样下保留模型的贪心解码拒绝响应,同时让其他提示词从原始高温分布中采样。在 7 个模型和 3 个基准数据集上、T=2.0 时,RGD 将贪心拒绝保持率从直接采样的 91.9% 提升至平均 98.3%,非拒绝请求的中位延迟仅增加 2.2-4.3%,并保持至少 98.1% 的贪心非拒绝样本走高温采样路径。

正文

View PDF HTML (experimental)

Abstract:Recent advances in truncation-based sampling have helped mitigate drawbacks of high-temperature sampling such as neural text degeneration, thereby enabling greater diversity without sacrificing coherence. However, increasing the entropy of the token probability distribution via high temperatures has also been shown to weaken the model's refusal response. Existing solutions for maintaining the refusal behavior of LLMs either replace the model's own refusal decision with a separate safety classifier or alter its output distribution for every prompt. To address this gap, we propose refusal-gated decoding (RGD): an efficient sequential decoding approach which preserves a model's greedy decoding refusal response at high temperatures and samples all other prompts from its exact direct high-temperature distribution, while incurring minimal additional latency. RGD runs a short greedy probe that reuses the prompt's KV cache and exits as soon as it becomes incompatible with a learned set of refusal prefixes; it returns the greedy response if the probe remains compatible and otherwise discards the probe and samples from the original prompt. Across seven models and three benchmark datasets at T=2.0, RGD raises greedy-refusal preservation from 91.9% under direct sampling to 98.3% on average while adding only 2.2-4.3% to the median per-request latency of non-refusals across temperatures. Unlike prompt-screening baselines which route many greedy non-refusals to greedy decoding, RGD keeps at least 98.1% of greedy non-refusals on unchanged high-temperature sampling, thereby preserving the model's natural high-temperature sampling behavior. We also propose a residual-stream variant of our method which lowers this latency overhead to at most 0.5% with comparable prompt routing accuracy. Our work shows that unlocking greater diversity via high-temperature sampling need not erode a model's refusal behavior.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2607.20791 [cs.AI]
  (or arXiv:2607.20791v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2607.20791

arXiv-issued DOI via DataCite

Submission history

From: Phillip Howard [view email]
[v1] Wed, 22 Jul 2026 23:33:51 UTC (206 KB)
[v2] Mon, 5 Oct 2026 22:35:13 UTC (353 KB)

来源:arXiv:cs.AI · arxiv.org