跳到正文
arXiv:cs.CL· Shiliang Xiao·· 3 小时前AI 评分17

TACS:面向 LLM 越狱后缀优化的轨迹感知候选选择框架(已被作者撤回)

TACS: Trajectory-Aware Candidate Selection for LLM Jailbreak Suffix Optimization

AI 导读

一篇提出 TACS 的 arXiv 论文已被作者 Shiliang Xiao 撤回,原因是理论分析存在错误,影响了主要结论的有效性。TACS 是面向 LLM 越狱后缀优化的轨迹感知候选选择框架,通过轨迹感知代理、参考策略正则化和判别器估计的卡方校正来稳定候选选择,在 HarmBench 上同等搜索预算下优于强基线并提升攻击成功率。

正文

This paper has been withdrawn by Shiliang Xiao

No PDF available, click to view other formats

Abstract:Gradient-based jailbreak suffix optimization methods typically update the suffix by retaining the candidate with the lowest current loss. We show that this seemingly natural design is fundamentally myopic: candidates that look better under the current-step proxy often fail to produce better jailbreak outcomes later in the search, revealing a form of selection-stage reward hacking. This suggests that candidate selection, rather than candidate generation alone, is a hidden bottleneck in suffix optimization. To address this issue, we propose TACS, a trajectory-aware candidate selection framework for jailbreak suffix optimization. Instead of selecting candidates solely by their immediate loss, TACS augments per-step evaluation with a trajectory-aware proxy and stabilizes selection with reference-policy regularization and a discriminator-estimated chi-squared correction, encouraging choices that remain effective beyond the current step. Experiments on HarmBench show that TACS consistently outperforms strong baselines under the same search budget, substantially improving attack success rates while exhibiting more stable optimization behavior throughout the search. Our findings highlight that mitigating selection-stage reward hacking caused by myopic candidate selection is critical for improving jailbreak suffix optimization.
Comments: We identified an error in the theoretical analysis, which affects the validity of the main conclusions of the manuscript. Since the current version does not adequately support this conclusion, we have decided to withdraw the paper
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2608.29564 [cs.CL]
  (or arXiv:2608.29564v3 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2608.29564

arXiv-issued DOI via DataCite

Submission history

From: Shiliang Xiao [view email]
[v1] Sun, 30 Aug 2026 05:18:28 UTC (1,906 KB)
[v2] Tue, 1 Sep 2026 03:40:32 UTC (1,906 KB)
[v3] Wed, 7 Oct 2026 16:21:07 UTC (1 KB) (withdrawn)

来源:arXiv:cs.CL · arxiv.org