跳到正文
arXiv:cs.LG· Jaturong Kongmanee, Smile Thanapattheerakul·· 7 小时前AI 评分39

Latent Diagnostic Taxonomy:面向提示注入检测的分类器构建与决策诊断框架

The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection

AI 导读

研究者提出 Latent Diagnostic Taxonomy 框架,用于构建安全防护层分类器并诊断其哪些高置信决策可信。该框架通过交叉验证选择嵌入维度、定位约 29% 训练样本的潜在支持向量,并按安全信任、启发式偏差与覆盖、上下文不足三类进行路由。在公开提示注入数据集上,约 77% 的高置信决策对删除单个 token 不稳健,且可分为置信度校准失败与可被利用的捷径两种模式。

正文

View PDF HTML (experimental)

Abstract:This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that identifies which of the classifier's confident decisions can be trusted. This framework, the Latent Diagnostic Taxonomy, consists of (i) constructing a dimensionality-optimized classifier, in which the embedding dimensionality is empirically selected via cross-validated performance rather than fixed a priori, (ii) locating a relatively small set of latent support vectors (~ 29% of total training examples) representing influential prompts for identifying tokens that alter the classifier's predicted labels, and (iii) utilizing such tokens and their associated attack magnitudes for constructing a diagnostic taxonomy. This diagnostic taxonomy provides an end-to-end guideline for flagging prompts that require different treatments: rely Safely on the classifier's decision; flag Heuristic Bias and Heuristic Override cases; route Insufficient Context cases for further human/safety review. Applying the framework to a classifier trained on a public prompt injection dataset, we find that a substantial fraction of its confident decisions (~ 77%) are not robust to removing a single token, and that this brittleness separates into two distinct failure patterns: a confidence calibration failure and a genuinely exploitable shortcut. For each zone of the taxonomy, we also recommend strategies for remediating diagnosed prompts. We illustrate the framework as a series of steps, demonstrating how each step operates.
Comments: 10 pages, 5 figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR)
Cite as: arXiv:2608.26423 [cs.LG]
  (or arXiv:2608.26423v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2608.26423

arXiv-issued DOI via DataCite

Submission history

From: Jaturong Kongmanee [view email]
[v1] Wed, 26 Aug 2026 21:55:15 UTC (233 KB)
[v2] Tue, 6 Oct 2026 15:27:36 UTC (235 KB)

来源:arXiv:cs.LG · arxiv.org