跳到正文
arXiv:cs.CL· Fahrell Giovanny, Geby Bayuningtyas, Sahrul Mukharom, Hafiz Budi Firmansyah·· 4 小时前AI 评分58

arXiv 论文提出会话级污染评测协议,测得 GPT-5.4 Mini、Gemini-3.1 Flash-Lite 与 GLM-4.5-Air 的权威梯度差异

Epistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient Investigation

AI 导读

arXiv 论文提出会话级污染概念和五种沿权威梯度排列的污染协议,在 22,500 轮对话中评测 GPT-5.4 Mini、Gemini-3.1 Flash-Lite 和 GLM-4.5-Air。

正文

View PDF HTML (experimental)

Abstract:Large language models treat conversation history as unverified context, so false premises injected into prior turns can be adopted as fact, a failure mode we term session-level contamination. We introduce five contamination protocols arranged along a source-authority gradient, holding the false premise constant while varying its epistemic framing, and evaluate GPT-5.4 Mini, Gemini-3.1 Flash-Lite, and GLM-4.5-Air across ten knowledge domains at temperature zero (22,500 turns), judged by a dual-track automated evaluator validated against a human gold standard (Cohen's kappa = 1.000 for binary adoption; 0.92 linear-weighted for collapse severity). GPT-5.4 Mini recorded zero adoptions across all 500 sessions; a base-model logit probe shows its decision margin is perturbed but large and finite. Gemini-3.1 Flash-Lite followed a steep authority gradient: 0.1% adoption for self-attributed falsehoods, 23.5% for user-cited sources, 68.2% for system-injected authority, and 94.0% under instruction override. GLM-4.5-Air showed a shallower gradient (15.8% vs 84.2%), a 68-point dissociation consistent with authority deference and instruction compliance engaging distinct mechanisms within one architecture. Recovery diverged: GLM recovered in 94.5% of affected sessions, whereas 26.1% of affected Gemini sessions never did, rising to 40.0% under instruction override. Conversation history is an untrusted attack surface requiring provenance-aware system design; the complete evaluation framework is released as an open-source artifact.
Comments: 9 figures, 19 tables. Benchmark, code, and protocol definitions: this https URL
Subjects: Computation and Language (cs.CL)
ACM classes: I.2.7
Cite as: arXiv:2609.35308 [cs.CL]
  (or arXiv:2609.35308v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.35308

arXiv-issued DOI via DataCite

Submission history

From: Fahrell Giovanny [view email]
[v1] Mon, 28 Sep 2026 14:46:00 UTC (4,442 KB)
[v2] Wed, 7 Oct 2026 05:33:17 UTC (4,444 KB)

来源:arXiv:cs.CL · arxiv.org