arXiv:cs.LG· Elad David, Max Fomin·· 3 小时前AI 评分40
分类后缀如何提升 LLM 智能体激活探针的恶意输入检测
Prompted to Discriminate: Generalizing Malicious-Input Probes in the Wild
AI 导读
在 13 个安全基准和 Llama-3.1-8B、Qwen3.5-9B、Gemma-4-12B 三个开源模型上,用户轮次后追加分类指令可让单点激活探针的分布外检测 AUC 最高提升约 4 点。增益来自分类格式本身而非具体标准,无内容标签的指令与真实恶意/良性标准效果相当。该收益可延续到生产环境使用的多点池化探针,但最佳后缀取决于模型与读出方式。
正文
Abstract:LLM agents increasingly rely on activation probes as runtime monitors for prompt injection, jailbreaks, and unsafe requests, reading the model's own hidden state to catch a harmful input before the agent acts on it. A cheap, increasingly common move, borrowed from LLM-as-judge prompting, is to append a short classification instruction after the user's turn and read the probe at that point, to sharpen it: the instruction asks the model to represent the incoming request as a class, concentrating the signal the probe must separate, at negligible serving cost. But does the wording of that suffix matter, and does its benefit hold in the wild, on attack types the probe never saw in training, the regime a deployed monitor faces? We test this with a controlled ladder of post-user suffixes under strict leave-one-dataset-out (LODO) evaluation across 13 safety benchmarks (jailbreak, injection, and benign chat) and three open-weight model families (Llama-3.1-8B, Qwen3.5-9B, Gemma-4-12B). On a single-position probe, a classification suffix consistently improves out-of-distribution detection over no suffix (up to ~4 AUC points); yet which suffix matters: prompting the model to classify the input, even into content-free labels, reliably wins; an off-topic or merely-attentive suffix helps little. The gain comes from the classification format, not the named criterion: a content-free suffix matches the real malicious/benign one, with the criterion adding precision only at strict thresholds. This is not an artifact of the single-position read: the benefit carries to the multi-position pooling probes used in production (attention, multi-max, MLP), though the best-performing suffix there is readout-dependent. Served through a KV-cache fork, it is a cheap drop-in for any activation-probe monitor, though not an automatic win: which suffix helps, and by how much, depends on the model and the readout.
| Comments: | 40th Conference on Neural Information Processing Systems (NeurIPS 2026). Workshop: Agents in the Wild: Safety, Security, and Beyond |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02413 [cs.LG] |
| (or arXiv:2610.02413v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02413 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Max Fomin [view email]
[v1]
Thu, 1 Oct 2026 19:36:57 UTC (1,631 KB)
来源:arXiv:cs.LG · arxiv.org