跳到正文
arXiv:cs.CL· Yangfan Hu, Xuhan Tong, Haoyue Bai, Xi Ding, Shashank Muralidhar Bharadwaj, Siyang Cao, Robert Nowak, Jiawei Zhang·· 4 小时前AI 评分44

语言模型为何产生幻觉:用推理对抗先验的测试研究

Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors

AI 导读

研究将大语言模型幻觉归因于"推理错位",即提示词支持的答案与统计上显著的潜在关联所偏好的答案不一致,并用潜键-任务模型证明预训练频率失衡会让捷径路径压过约束敏感路径。作者提出 TrapQA 诊断测试集,包含 ScientistQA 和 Real-Life Constrained QA 两部分,结果显示幻觉可源于有偏的潜在推理,而非单纯知识缺失。

正文

View PDF HTML (experimental)

Abstract:Large language models often produce hallucinated answers that violate prompt-level constraints. A key diagnostic question is whether these failures reflect missing knowledge, or whether the model has the relevant information but follows the wrong inference path. We study this phenomenon as inference misalignment: a mismatch between the answer supported by the prompt and the answer favored by statistically salient latent associations. We formalize this view with a latent key-task model, in which pretraining-frequency imbalance can cause a shortcut path to dominate the constraint-sensitive path and induce positive inference loss. The framework predicts two failure modes: task-retrieval bias in entity disambiguation and key-selection bias in action choice. We introduce TrapQA, a controlled diagnostic testbed with two components. ScientistQA tests disambiguation among similar scientists with supplementary factual probes, while Real-Life Constrained QA tests everyday constraint following under salient shortcuts. Our results show that hallucination can arise from biased latent inference rather than absent knowledge alone.
Comments: Findings of EMNLP 2026
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2607.00447 [cs.CL]
  (or arXiv:2607.00447v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2607.00447

arXiv-issued DOI via DataCite

Submission history

From: Yangfan Hu [view email]
[v1] Wed, 1 Jul 2026 05:02:43 UTC (1,233 KB)
[v2] Thu, 1 Oct 2026 22:15:47 UTC (1,241 KB)

来源:arXiv:cs.CL · arxiv.org