跳到正文
arXiv:cs.LG· Boyuan Chen, Yehia Dawoud, Hailemariam Mersha, Minghao Shao, Siddharth Garg, Ramesh Karri, Muhammad Shafique·· 4 小时前AI 评分33

图像哪一属性承载了越狱?对图生文越狱的受控剖析

Which Image Property Carries the Jailbreak? A Controlled Dissection of Image-to-Text Jailbreaks

AI 导读

一项受控研究在 313 条 StrongREJECT 切片上考察了四类已发表攻击的图像侧因素,使用五个多模态模型并额外评估 InternVL3.5-8B。结果显示纯有害查询加不加无关良性图像都很难攻击成功,攻击图像则显著提升成功率;逐 tile 熵与 JPEG 大小无法可靠区分攻击 tile 与尺寸匹配的良性干扰图。

正文

View PDF HTML (experimental)

Abstract:Image-to-text jailbreaks place harmful intent in text, image content, or the relationship between them. We examine image-side factors across four published attack families on a 313-prompt StrongREJECT slice, using five multimodal models and an additional appendix evaluation of InternVL3.5-8B. The harmful instruction is held constant across conditions; the baseline matrix uses one draw per prompt, and paired ablations use three draws with an automated rubric judge. A bare harmful query, with or without a benign unrelated image, produces little attack success on most victims, while attack images substantially increase it. First, per-tile entropy and JPEG size do not reliably distinguish attack tiles from size-matched benign distractors, limiting density-only screening. Second, earlier tile-count ladders were confounded by payload visibility. A corrected region-count test found no detectable effect, so the role of tile-count structure remains unresolved. Third, on Qwen3-VL-8B, the E4 manipulation that removes query-specific relatedness lowers ASR by about 0.12. This supports a bounded attribution to the relatedness manipulation, although image-text congruence remains unmeasured. A within-category control reproduces the direction on overlapping stimuli. Five tests survive the global statistical correction, but only E4 supports attribution to one measured descriptor; the within-category result is a robustness check, not a separate attribution. These conclusions remain conditional on the rubric judge.
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.07009 [cs.CR]
  (or arXiv:2610.07009v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.07009

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Boyuan Chen [view email]
[v1] Sun, 4 Oct 2026 13:23:31 UTC (1,022 KB)

来源:arXiv:cs.LG · arxiv.org