跳到正文
arXiv:cs.LG· Yibo Zhang, Tianrong Guan, Liang Lin, Puze Wang, Jin Wang, Qingsong Wen·· 4 小时前AI 评分54

arXiv 论文提出多轮对话中的答案侧后门攻击,触发词由模型自生成

The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models

AI 导读

arXiv 论文(arXiv:2610.07723)提出针对多轮对话 LLM 的答案侧后门攻击:攻击者用良性首轮提示诱导模型自己生成一个看似无害的词,该词进入对话历史后成为触发器,后续有害查询到来时模型绕过安全拒绝,而用户输入始终干净。

正文

View PDF HTML (experimental)

Abstract:Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word. Once merged into the dialogue history, this self-generated word becomes the trigger. When a later harmful query arrives, the model detects its own trigger and bypasses its safety refusal, while the user input stays perfectly clean. Across four LLMs, our attack reaches near-perfect Attack Success Rates, approaching 100\% at only a 5\% poisoning rate, while preserving general utility and clean-input safety, and it evades mainstream input-centric defenses. Representation-level analysis shows that the self-generated trigger consistently suppresses the model's refusal signal, exposing a critical blind spot in current LLM defenses.
Subjects: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
Cite as: arXiv:2610.07723 [cs.CR]
  (or arXiv:2610.07723v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.07723

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yibo Zhang [view email]
[v1] Tue, 6 Oct 2026 04:17:18 UTC (2,242 KB)

来源:arXiv:cs.LG · arxiv.org