跳到正文
arXiv:cs.LG· Valeria Ruscio, Keiran Thompson·· 5 小时前AI 评分49

LLM 为何会幻觉出自己能解码的答案:解码与选择之间的差距

On the Tip of the Tongue: Why LLMs Hallucinate Answers They Can Decode

AI 导读

研究区分了 LLM 首答案 token 的"读取"与"写入":在三种 reader 和随机标签控制下,相当一部分失败答案仍可从中间残差状态解码,但最终读出层却选中了另一个内容 token。

正文

View PDF HTML (experimental)

Abstract:A language model can give the wrong answer even when the correct answer is decodable from its intermediate states. To study this gap between decodability and selection, we distinguish \textit{read} from \textit{write} at the first answer token. Read asks whether the gold token can be decoded from intermediate residual states under same-relation decoy controls. Write asks whether the final readout ranks that token first among content tokens. Under three different readers, with a randomized-label control, a substantial fraction of failures remain readable while another content token is selected. We explain this through the selection margin at the final readout, the difference between the answer logit and the logit of its strongest alternative, which is answer support minus alternative support, and can also be split into a context-averaged baseline linked to token frequency and an item-specific term. Setting the answer support to the level typical of successful generations is sufficient to recover first-token selection for the majority of failures in most of the models we study; the original alternative remains ahead in most remaining failures under this edit, and this outcome follows directly from the readout geometry. Removing the frequency direction alone shifts selection but rarely recovers the answer. Prompt variants of the same fact that succeed supply support that transfers to failing variants through the residual stream and through late MLP outputs, with less consistent effects through late attention. First-token recovery leaves most full answers wrong, which limits the recovery achieved by these edits and separates three things that are easily conflated, decodability, recoverability, and generation.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2603.13911 [cs.AI]
  (or arXiv:2603.13911v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2603.13911

arXiv-issued DOI via DataCite

Submission history

From: Valeria Ruscio [view email]
[v1] Sat, 14 Mar 2026 11:55:55 UTC (30,840 KB)
[v2] Fri, 2 Oct 2026 16:57:04 UTC (45 KB)

来源:arXiv:cs.LG · arxiv.org