跳到正文
arXiv:cs.AI· Lin Tian, Marian-Andrei Rizoiu·· 4 小时前AI 评分49

语言模型虚假信息无需触发词:从事实回答到下游决策的审计缺口

Misinformation Without Triggers: From Factual Answers to Downstream Decisions

AI 导读

研究发现语言模型即使直接事实回答正确,下游决策仍可能被虚假训练数据误导,形成"审计缺口"。在 Guess the Capital 任务中,八个模型在剂量 1,000 时直接注入选择率达 95.8%–100%,但游戏选择仅比真实训练增加 1.7–14.4 个百分点;通过直接探测的事实仍会将决策推向注入答案。

正文

View PDF HTML (experimental)

Abstract:Language models learn from web documents, some of them false, and false content can reach a model's answer to a factual question and the summaries and decisions that use it. Most data-poisoning studies add a trigger to the training data and activate it in the prompt. False documents can also change factual responses without any trigger, but we do not know whether the direct answer predicts the decision. In this work, we follow false content past the answer and find an \emph{audit gap} between what a direct probe reports and what the model then does, comparing false training with matched truthful controls in a controlled decision task, \emph{Guess the Capital}, where a fixed decoder turns factual answers into a scored card choice, and on a misleading claim from Facebook posts about the 2019--20 Australian bushfires. Across eight models at dose 1,000, direct injected-choice rates reach 95.8--100\%, while injected game choices increase by 1.7--14.4 percentage points over matched truthful training. The gap runs the other way too. Facts that pass the direct probe still push decisions toward the injected answer, and game accuracy drops further than those choices explain. Truthful correction brings the fact back but not the decisions built on it. We then look into the real-world bushfire case, models trained on the false posts say that people were arrested for arson even when they lose the inflated count, and in a count-by-wording factorial the misleading arrest wording produces arrest assertions even when the training count stays at 24. In simpler terms, \textbf{a correct factual answer does not guarantee a correct decision, and losing the injected number does not remove the misleading story}.
Comments: 35 pages, 11 figures, 16 tables
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02886 [cs.CL]
  (or arXiv:2610.02886v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.02886

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Lin Tian [view email]
[v1] Fri, 2 Oct 2026 06:24:54 UTC (1,237 KB)

来源:arXiv:cs.AI · arxiv.org