arXiv:cs.LG· Toluwani Aremu, Manit Baser, Mohan Gurusamy, Nils Lukas, Dinil Mon Divakaran·· 5 小时前AI 评分45
开放权重 LLM 中用于滥用检测的触发标签机制为何脆弱
The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs
AI 导读
研究形式化了开放权重 LLM 中的触发标签机制,将其分为解码时注入信号的 token 级触发标签与学习后门式关联的权重级触发标签,并提出统一攻击框架 Untag。以钓鱼内容为案例的评估显示,现有触发标签机制在攻击下完全失效,因此不应被视为鲁棒的滥用检测器。
正文
Abstract:Open-weight language models can be downloaded, modified, and deployed beyond their developers' control, limiting the effectiveness of centrally enforced safeguards. Recent work has therefore proposed \emph{trigger-tag} mechanisms that produce a detectable signal when a model is used under a target condition, such as generating phishing contents. Although these mechanisms borrow from established techniques, their use for conditional misuse detection in open-weight LLMs is relatively new. Therefore, existing research works have not systematically studied the robustness of trigger-tag mechanisms under adversarial attacks. To close this gap, (i)~we formalize trigger-tags and distinguish \emph{token-level trigger-tags}, which introduce watermark-inspired signals during decoding, from \emph{weight-level trigger-tags}, which learn backdoor-inspired associations between target conditions and detectable model behavior. Furthermore, (ii)~we introduce \Untag, a unified attack framework that organizes their mechanism-specific attack surfaces into a common taxonomy. We evaluate representative token-level and weight-level trigger-tags using phishing as a case study. We find that while trigger-tags may provide useful evidence in controlled settings, our attacks render the existing trigger-tag mechanisms to be entirely ineffective. Consequently, we argue that these mechanisms should not be treated as robust misuse detectors when attackers can transform outputs or modify open weights.
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.03124 [cs.CR] |
| (or arXiv:2610.03124v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03124 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Toluwani Aremu [view email]
[v1]
Fri, 2 Oct 2026 10:45:14 UTC (575 KB)
来源:arXiv:cs.LG · arxiv.org