arXiv:cs.AI· Pranjal Garg·· 6 小时前AI 评分45
可解释性探针中的 0.6 有多高?下限、上限与余量
How High Is 0.6? Floors, Ceilings, and Headroom in Interpretability Probing
AI 导读
研究者提出用「下限」与「上限」两个参考点解读探针分数,二者之间的差距称为余量,即探针能证明模型算出简单输入之外内容的范围。在上下文元分析任务中,分布偏移下探针分数下降、预测误差上升 12–15 倍,但模型仍恢复了相近比例的余量,说明是数据丢失信息而非表征问题;对 scGPT 及四项 LLM 探针研究的复检显示,部分结论仅凭输入文本即可解释。
正文
Abstract:Probes are the workhorse of interpretability. If a model's hidden states predict a variable, the model is said to represent it. But a probe score has no fixed meaning. An $R^2$ of 0.6 may only reflect what the input already gives away, and the same score can mean different things on different data. We propose reading every probe score against two reference points: a floor, what a declared set of simple inputs already predicts, and a ceiling, what the full input can predict. The gap between them, the headroom, is the range in which a probe can show that a model computes something beyond the simple inputs. We prove that headroom vanishes in two ways: the target stops depending on a hidden variable the model must infer, or the input stops revealing it. We test this on transformers trained for in-context meta-analysis, which must infer the hidden heterogeneity between studies to weight them correctly, and where both reference points are known. Under distribution shift, probe scores fall and prediction error rises $12$--$15\times$, yet the model recovers a similar share of the headroom, indicating that the data lost information, not the representation. We then analyze the real models. The single-cell foundation model scGPT encodes biological variability only partially. We also revisit four influential LLM probing studies, which claim that models represent geography, the state of an Othello board, truth, and the demographics of their users. Against a floor computed from the input text alone, some of these claims hold, while others are largely explained by the text itself.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08544 [cs.CL] |
| (or arXiv:2610.08544v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08544 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pranjal Garg [view email]
[v1]
Tue, 6 Oct 2026 15:32:25 UTC (1,020 KB)
来源:arXiv:cs.AI · arxiv.org