arXiv:cs.LG· Fengwei Tian, Ravi Tandon·· 3 小时前AI 评分41
意图隐藏越狱:组合攻击的信息论框架
Intent-Hiding Jailbreaks: An Information-Theoretic Framework for Compositional Attacks
AI 导读
研究从信息论视角分析组合式意图隐藏越狱,提出"先验-后验匹配"方法:通过选择辅助任务使含有害目标的任务束平均有害概率与整体先验一致,从而隐藏意图。在查询无关设定下,精确匹配在任务束大小约束下是计算困难的,作者给出分数权重的最优注水解,并刻画满足安全阈值的最小任务束。在多个开源模型上评估显示,组合查询在给定搜索预算下可诱发超出直接请求基线的目标行为,但任务束增大时响应级目标保持度趋于下降。
正文
Abstract:Recent work has shown that large language models (LLMs) can be vulnerable to jailbreak attacks in which harmful intent is obscured through composition with benign tasks. A harmful request refused in isolation may elicit a different response when embedded within a larger, seemingly benign query. We study these compositional intent-hiding jailbreaks from an information-theoretic perspective. Our formulation associates each task with an estimated probability of being judged harmful: the average over the full task collection defines the prior probability of harmful intent, while the average over a selected bundle containing the target defines the posterior. Selecting auxiliary tasks so that these averages agree, which we call prior-posterior matching, leaves the estimated intent unchanged even though the harmful target remains in the bundle.
We study two settings that differ in whether query construction is part of the optimization. In the query-independent setting, tasks are selected without regard to how they will be expressed in the final query. We show that exact prior-posterior matching under a bundle-size constraint is computationally hard, derive an optimal water-filling solution for fractional weights, and characterize the smallest bundle satisfying a prescribed safety threshold. In the query-dependent setting, task selection and query construction are considered jointly, and intent concealment and target preservation are evaluated on the resulting query. We evaluate jailbreak effectiveness and preservation of the target behavior across bundle sizes, query generators, and several open-source models. These results show that compositional queries can elicit target behaviors beyond the direct-request baseline under the evaluated search budgets, while revealing a trade-off: as bundle size increases, response-level target preservation tends to decrease for several models.
| Subjects: | Cryptography and Security (cs.CR); Information Theory (cs.IT); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02302 [cs.CR] |
| (or arXiv:2610.02302v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02302 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Fengwei Tian [view email]
[v1]
Thu, 1 Oct 2026 17:57:04 UTC (391 KB)
来源:arXiv:cs.LG · arxiv.org