arXiv:cs.CL· Orion Reblitz-Richardson·· 4 小时前AI 评分43
LLM 如何组织与建构道德知识:一项基于线性探针的研究
How Language Models Organize and Structure Moral Knowledge
AI 导读
研究者用六个独立线性探针在开放权重语言模型上分别探测道德基础理论(MFT)的六类道德维度,发现这些方向既不塌缩为单一道德检测器也不彼此孤立,而是张成接近最大数量的独立维度并共享一个正向公共分量,其平均成对余弦相似度为 0.26,而匹配的非道德概念组仅为 0.013。
正文
Abstract:How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically.
We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component. The shared component is the signature of integration, and it is moral-specific relative to a matched non-moral concept battery built identically (mean pairwise cosine 0.26 vs. 0.013).
The geometry is consistent across architectures and scale and reaches its integration regime early in pre-training, well before probe accuracy saturates. The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 10 distinct splits exist, so it cannot reject at the 0.05 level) but rather reflects corpus statistics. Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched-pair baseline, while the majority of its variance encodes conflict-specific structure. The model represents moral tension itself, not a pre-resolved judgment.
| Comments: | 32 pages, 16 figures. Code and outputs at this https URL |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| ACM classes: | I.2.6; I.2.7 |
| Cite as: | arXiv:2608.27402 [cs.CL] |
| (or arXiv:2608.27402v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2608.27402 arXiv-issued DOI via DataCite |
Submission history
From: Orion Reblitz-Richardson [view email]
[v1]
Thu, 27 Aug 2026 17:30:30 UTC (221 KB)
[v2]
Wed, 7 Oct 2026 17:49:26 UTC (222 KB)
来源:arXiv:cs.CL · arxiv.org