arXiv:cs.AI· Luman Zhao, Minghui Xu, Yue Zhang, Yijun Yang·· 3 小时前
LTBD:用于提示词注入防御的可学习信任边界分隔符
LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense
AI 导读
针对 LLM 易受提示词注入攻击的问题,研究者提出轻量防御方法 LTBD,在不改动 LLM 参数的前提下,用少量可学习分隔符在输入中显式编码信任边界,区分可信用户指令与不可信外部数据。LTBD 在 AlpacaFarm 上实现 0.00% ASR,在 TaskTracker 上仅 0.11-0.19% ASR,并在攻击者完全了解防御的自适应攻击下依然有效,且推理开销可忽略。
正文
Abstract:Large language models (LLMs) perform remarkably well on complex tasks, yet remain highly vulnerable to prompt injection attacks, where malicious instructions embedded in external data can override user intent. Existing defenses remain limited by model fine-tuning requirements, vulnerability to adaptive attacks, or reliance on brittle handcrafted prompts. We argue that a fundamental source of this vulnerability is the lack of an explicit representation of trust provenance. To address this, we introduce Learnable Trust-Boundary Delimiters (LTBD), a lightweight defense that explicitly encodes trust boundaries in the input while keeping the LLM parameters unchanged. LTBD uses a small number of learnable delimiters to distinguish trusted user instructions from untrusted external data, enabling the model to better respect the intended trust hierarchy. Experimental results show that LTBD substantially outperforms inference-time defenses and performs competitively with training-based approaches, while preserving benign-task utility and introducing negligible inference overhead. In particular, LTBD achieves 0.00% ASR on AlpacaFarm and only 0.11-0.19% ASR on TaskTracker. LTBD also remains effective under adaptive attacks, where adversaries have full knowledge of the defense and explicitly attempt to bypass it.
| Comments: | 5 pages, 3 tables, 1 figures |
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11634 [cs.CR] |
| (or arXiv:2610.11634v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11634 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Luman Zhao [view email]
[v1]
Thu, 8 Oct 2026 10:13:24 UTC (210 KB)
来源:arXiv:cs.AI · arxiv.org