针对黑盒 LLM 的受控解码攻击框架
Controlled Decoding Attacks on Black-Box LLMs
研究者提出一种针对黑盒 LLM 的越狱框架,仅通过纯文本续写接口即可绕过安全对齐,无需访问模型权重或 token 概率。该框架包含基于采样的分布重建、风险门控残差控制和推测式多 token 执行三个组件,将采样成本集中在少数关键位置。在四个目标端点和三个基准测试中,其平均得分在多数对比中优于基线方法。
Published on Sep 29
Authors:
,
,
,
,
,
Abstract
Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak trajectories are concentrated at a small subset of positions, motivating selective control. We introduce , a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation. Sample-Based Distribution Reconstruction combines sampled outputs with a prior over unobserved actions to obtain a usable control signal. Risk-Gated Residual Control uses the evolving response prefix to decide when to reconstruct and modify the distribution, concentrating sampling costs at selected positions. Speculative Multi-Token Execution further amortizes target calls by verifying and accepting draft prefixes that require no intervention. Across four target endpoints and three benchmarks, achieves the highest mean score most comparisons against baselines.
View arXiv page View PDF Add to collection
Get this paper in your agent:
hf papers read 2609.36956
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2609.36956 in a model README.md to link it from this page.
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2609.36956 in a dataset README.md to link it from this page.
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2609.36956 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.
来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co