跳到正文
arXiv:cs.LG· Kaley Brauer, Claudio Mayrink Verdun, Samuel Marks·· 2 天前AI 评分67

arXiv 论文:从填充 token 中解码大语言模型的隐藏计算

Reading Between the Dots: Decoding Hidden Computation across Filler Tokens

AI 导读

Kaley Brauer、Claudio Mayrink Verdun、Samuel Marks 等人的论文(arXiv:2607.03502,已被 NeurIPS 2026 接收)研究前沿大模型在点状填充 token 上的多步推理。

正文

View PDF HTML (experimental)

Abstract:Frontier LLMs can perform multi-step reasoning over content-free filler tokens like dots or counting sequences, producing correct answers with no visible chain-of-thought (CoT). This is a limit case for behavioral oversight, where surface tokens carry no information about the underlying reasoning. But hidden from the output is not the same as hidden from us. On four task families (fact retrieval, parallel numeric composition, string manipulation, and in-context computation), two open-weights frontier models (DeepSeek V3, Kimi K2) compute over filler tokens in a legible way: attention routes the question through the filler region to the answer, logit-lens readouts show retrieved facts emerging early and their composition crystallizing in late layers, and KV-cache transplants at filler positions causally swap outputs between examples. We introduce an unsupervised decoding pipeline that takes only hidden states as input and recovers intermediate values with 82-94% accuracy (best LLM judge) across both models and all four tasks, without ground-truth labels or training. Even without a judge, the hidden values are already directly in the pipeline's top-2 tokens 35-85% of the time. The uplift persists whether the filler is prefilled or the model generates the filler itself. On these cleanly decomposable tasks, hidden computation that defeats behavioral CoT monitoring is readable from the residual stream, which suggests that monitorability is a property of the model's full computational trace rather than only its surface tokens.
Comments: Accepted to NeurIPS 2026, 10 main paper pages, 27 appendix pages
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2607.03502 [cs.CL]
  (or arXiv:2607.03502v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2607.03502

arXiv-issued DOI via DataCite

Submission history

From: Kaley Brauer [view email]
[v1] Fri, 3 Jul 2026 17:18:34 UTC (2,596 KB)
[v2] Thu, 1 Oct 2026 04:10:02 UTC (2,579 KB)

来源:arXiv:cs.LG · arxiv.org