跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Huawei Lin, Yingjie Lao, Tony Geng, Tan Yu, Weijie Zhao·· 15 小时前AI 评分42

UniGuardian:首个免训练 LLM 检测器,统一防御提示词注入、后门与对抗攻击

UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models

AI 导读

UniGuardian 是首个免训练的 LLM 检测器,无需预知攻击类型即可同时检测成功激活的提示词注入、后门和对抗攻击,其共享机制通过度量结构化提示词扰动对模型输出分布的影响来识别隐藏触发器。

正文

View PDF HTML (experimental)

Abstract:Large Language Models (LLMs) are vulnerable to attacks like prompt injection, backdoor attacks, and adversarial attacks, which manipulate prompts or models to generate harmful outputs. In this paper, departing from traditional deep learning attack paradigms, we explore their intrinsic relationship and collectively term them Prompt Trigger Attacks (PTA). This raises a key question: Given a prompt, can we tell whether a hidden trigger is steering the model's behavior? We propose UniGuardian, to the best of our knowledge the first training-free LLM detector to jointly detect successfully activated prompt injection, backdoor, and adversarial attacks without knowing the attack type. Its shared mechanism measures how structured prompt perturbations shift the model's output distribution. Additionally, we introduce a single-forward strategy to optimize the detection pipeline, enabling simultaneous attack detection and text generation within a shared batched forward pass at each decoding step. Our experiments confirm that UniGuardian accurately and efficiently identifies trigger-activated prompts in LLMs.
Comments: 25 Pages, 13 Figures, 11 Tables. Accepted to Findings of AACL-IJCNLP 2026. Keywords: Attack Defending, Security, Prompt Injection, Backdoor Attacks, Adversarial Attacks, Prompt Trigger Attacks
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2502.13141 [cs.CL]
  (or arXiv:2502.13141v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2502.13141

arXiv-issued DOI via DataCite

Submission history

From: Huawei Lin [view email]
[v1] Tue, 18 Feb 2025 18:59:00 UTC (562 KB)
[v2] Thu, 1 Oct 2026 17:30:02 UTC (2,653 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org