跳到正文
arXiv:cs.AI· Ritvij Sharma, Russell Dlugosz, Ryan Zhou, Maheep Chaudhary·· 6 小时前AI 评分42

面向 LLM 的纵深防御:评估记忆门控对激活诱导与记忆诱导谄媚的防御效果

Defense-in-Depth for LLMs: Evaluating Memory Gates Against Activation-Induced and Memory-Induced Sycophancy

AI 导读

研究人员提出一个 2×2 纵深防御框架,将内部激活引导与外部记忆处理分离,并在 MemSyco-Bench(全部 1,550 题,防御条件在固定 250 题子样本上评判)上用三个 LLM 评委评估四个开源权重模型。

正文

View PDF HTML (experimental)

Abstract:Long-term memory allows Large Language Models (LLMs) to maintain personalized context across interactions, but retrieved user history can induce memory-induced sycophancy, causing models to favor stored user beliefs over objective evidence. Existing defenses primarily operate on retrieved context and are rarely evaluated jointly with internal behavioral bias. We introduce a $2 \times 2$ defense-in-depth framework separating internal activation steering from external memory handling. We extract sycophancy steering directions from 100 paired prompts and evaluate four open-weight models across 10 steering coefficients and five memory-defense configurations on MemSyco-Bench (answers for all 1,550 items; defense conditions judged on a fixed 250-item subsample), with three LLM judges. Three of the five configurations are new (rewriting every memory, a Router Gate that keeps, rewrites, or drops each memory, and dropping all memory); the other two are MemSyco's baselines. Selective Router Gate filtering preserves substantially more of MemSyco's average accuracy than complete memory removal, and this separation persists when the models are steered toward sycophancy. On Llama 3.1 8B with Router Gate, mild inverse steering ($\alpha = -1.5$) lowers judge-averaged sycophancy from 35.80% to 31.32% while average accuracy moves from 43.99% to 43.31%; this reduction has the same direction under all three judges but is not statistically significant (paired $p = 0.08$ to $0.63$ on 149 items). External memory filtering is the part of the design that holds up; our data do not show that inverse steering adds to it.
Comments: Accepted to NeurIPS (IAB, RTCA, AIWILD, and CL4FM)
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07403 [cs.AI]
  (or arXiv:2610.07403v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07403

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ritvij Sharma Mr. [view email]
[v1] Mon, 5 Oct 2026 21:16:06 UTC (83 KB)

来源:arXiv:cs.AI · arxiv.org