arXiv:cs.LG· Elan Barenholtz·· 6 小时前AI 评分63
arXiv 论文:静态词向量可复现 LLM 世界属性解码结果,质疑世界模型解读
World Properties without World Models: Distributional Associations and the Interpretation of Decoding Results from Language Models
AI 导读
Elan Barenholtz 在 arXiv 论文(arXiv:2603.04317)中指出,静态词嵌入在四个已发表案例(地点、时间、疼痛、情绪)上支持与 LLM 相当的线性解码:预测坐标和去世年份(R^2 = 0.42-0.59),区分疼痛句(held-out AUC 0.85-0.88),并对十二种情绪分类(AUC 0.84-0.88)。
正文
Abstract:A growing literature shows that variables can be linearly decoded from the activations of large language models (LLMs). These range from properties of the world, such as the locations of cities and the lifetimes of historical figures, to emotions and pain. Such findings are often taken as evidence that language models go beyond surface text statistics and form internal models of the world. We show that static word embeddings (fixed, context-insensitive representations learned from corpus statistics) of the same or matched stimuli support much of the same decoding. Across four published cases (place, time, pain and emotion), static vectors predict coordinates and year of death (R^2 = 0.42-0.59), separate pain from matched control sentences (held-out AUC 0.85-0.88), and classify twelve emotions in stories written to avoid naming them (AUC 0.84-0.88). Because static embeddings assign each word a single, context-independent vector, these results are a lower bound on what word associations alone can support. The LLMs retain clear advantages on representational tests, and causal and behavioral findings remain outside the scope of the baseline. On the original authors' entities, where we reproduce their Llama-2 results, the transformer's advantage lies mostly in placing historical figures in the right century and places in the right country, coarse sorting that richer word associations would be expected to improve; within those groups every representation orders items poorly. Static vectors for disambiguated Wikipedia entities, which carry the associations of a particular place or person rather than of the words in its name, close most of the remaining gap, matching Pythia-2.8B on coordinates and Llama-2-7B on year of death. These results indicate that decodability alone cannot distinguish a representation of a property from information already available in fixed distributional associations.
| Comments: | 22 pages, 3 figures, 10 tables. Substantially revised to include analyses of full released Gurnee & Tegmark datasets with Llama-2 and Pythia comparisons; replaces the earlier 100-city, 194-figure analysis; also includes entity-level vectors, and pain and emotion decoding |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2603.04317 [cs.CL] |
| (or arXiv:2603.04317v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2603.04317 arXiv-issued DOI via DataCite |
Submission history
From: Elan Barenholtz [view email]
[v1]
Wed, 4 Mar 2026 17:37:05 UTC (1,096 KB)
[v2]
Tue, 6 Oct 2026 05:43:50 UTC (1,749 KB)
来源:arXiv:cs.LG · arxiv.org