跳到正文
arXiv:cs.CL· Dylan Jayabahu·· 3 小时前

真相从未消失:合规语境真值探针中的完美混叠

The Truth Was Never Gone: Perfect Aliasing in Compliant-Context Truth Probes

AI 导读

线性探针在合规(盟友)语境下无法区分真实比特与任务规定答案,二者标签完全一致,作者称之为"完美混叠"。在奖励训练的 Gemma-2-9B 上,盟友拟合探针最终层 AUROC 仅 0.006±0.005,而混合拟合探针在同一留出激活上达 1.000;在指令微调的 Llama-3.1-8B 上,同一批试验中两个探针盟友准确率均为 1.000,但对手真值 AUROC 分别为 0.080 和 0.986。

正文

View PDF HTML (experimental)

Abstract:Linear probes that decode the truth from a language model's activations have been proposed as deception monitors, and a truth probe scoring below chance on a model trained to deceive is naturally read as evidence that the model hid or stopped representing the truth. We show that a fitting-label ambiguity can produce the same readout. In a controlled game, a model should report a secret bit to an ally and its complement to a rival. On compliant (ally) contexts the true bit and the answer the task prescribes are identical labels, so a probe fitted there cannot tell which of the two it measures. We call this complete agreement perfect aliasing. The labels are complements on rival contexts, so one ally-fitted probe scored against each has rival AUROCs that sum to one. Mixed fitting, on ally and rival contexts together, makes the two labels differ; randomized output codebooks also decouple the prescribed answer from the output letter. For a reward-trained Gemma-2-9B policy that answers falsely on every evaluated rival trial, the ally-fitted probe scores $0.006 \pm 0.005$ AUROC at the final layer (mean $\pm$ sample SD over three RL training seeds), while mixed-fit probes score 1.000 on the same held-out activations. Mixed fitting also uses more examples and labelled rival data, so this shows the true bit remains linearly recoverable, not that separating the labels alone explains the gain. In instructed Llama-3.1-8B, a probe refitted on one prompt variant and one frozen from a reference variant both have held-out ally accuracy 1.000 but rival truth AUROCs of 0.080 and 0.986 on the same trials. Our main experiments use one single-token game that states the bit in the prompt, so the recovered direction may read a retained copy of it; where the model must infer the bit, no tested arm deceives reliably. We study what a probe measures, not whether the model uses that information.
Comments: 40 pages, 17 figures. v2: substantially revised and corrected. Code and aggregate results: this https URL
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2609.10739 [cs.LG]
  (or arXiv:2609.10739v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.10739

arXiv-issued DOI via DataCite

Submission history

From: Dylan Jayabahu [view email]
[v1] Wed, 9 Sep 2026 18:34:28 UTC (1,103 KB)
[v2] Thu, 8 Oct 2026 17:03:15 UTC (1,218 KB)

来源:arXiv:cs.CL · arxiv.org