跳到正文
arXiv:cs.LG· Niveen O. Jaffal, Ahmet Yuksel, David Mohaisen·· 7 小时前AI 评分38

RAG-PIBench:面向可信 RAG 系统提示注入检测的防泄漏基准

RAG-PIBench: A Leakage-Aware Benchmark for Prompt-Injection Detection in Trustworthy RAG Systems

AI 导读

研究人员推出 RAG-PIBench,一个面向 RAG 提示注入检测的基准,包含 4,876 条上下文样本并划分为冻结的训练、验证与受保护测试集。在防泄漏构建流程与严格评测协议下,DistilBERT 取得最佳受保护测试表现(F1 = 0.896,PR-AUC = 0.968),TF-IDF SVM 与逻辑回归同样具备竞争力。

正文

View PDF HTML (experimental)

Abstract:Retrieval-Augmented Generation (RAG) systems are vulnerable to prompt-injection attacks embedded in retrieved content. We introduce RAG-PIBench, a benchmark for RAG-style prompt-injection detection containing 4,876 contextual examples across frozen train, validation, and protected-test splits. Using a leakage-aware construction pipeline and strict evaluation protocol, we compare keyword-based, semantic-reference, TF-IDF, and transformer-based detectors. DistilBERT achieves the best protected-test performance (F1 = 0.896, PR-AUC = 0.968), while TF-IDF SVM and logistic regression remain competitive. Our results demonstrate the value of leakage-aware benchmark design and strong sparse baselines for reliable prompt-injection detection in RAG systems.
Comments: 19 pages, 3 figures, 8 tables
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.08571 [cs.CR]
  (or arXiv:2610.08571v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.08571

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: David Mohaisen [view email]
[v1] Tue, 6 Oct 2026 15:47:47 UTC (162 KB)

来源:arXiv:cs.LG · arxiv.org