arXiv:cs.CL· Or Shafran, Mor Geva·· 5 小时前AI 评分40
容量、响应性与对齐:什么让语言模型的潜结构可被操控
Capacity, Responsiveness and Alignment: What Makes a Latent Structure Actionable
AI 导读
研究将语言模型潜结构的因果影响力拆解为容量、响应性与对齐三个因素,在 4 个模型家族、50 个概念上验证三者必须同时高才能有效干预:容量和响应性偏低分别使因果效果下降 84% 和 95%,对齐偏低则会反转效果、抑制概念表达。因果有效方向在不同上下文中构成变化的低维子空间,将线性探针训练限制在该子空间后,跨模型操控效果提升 17%-118%,概念检测仅下降 3%。
正文
Abstract:Localizing latent structures in the activation space of language models (LMs) is central to understanding and controlling their behavior. Yet, localized structures can differ substantially in their causal influence, raising the question of what makes a structure actionable. We tackle this question by casting causal influence as a product of three factors and showing empirically that they act as interpretable, distinct constraints: capacity, measuring the sensitivity of the model's output to movement along the structure, responsiveness, capturing how promotable the concept is given the current context, and alignment, reflecting how well the structure aligns with the context-specific representation of the concept. Across 4 LM families and 50 concepts, we observe that causal effectiveness requires all factors to be high; low capacity and responsiveness reduce it by 84% and 95%, respectively, while low alignment can reverse it, suppressing concept expression. Moreover, we find that causality is context-dependent rather than an intrinsic property of the structure, with causally effective directions forming a low-dimensional subspace that varies across contexts. By restricting the training of linear probes to this subspace, we introduce causal probes that achieve 17%-118% improvement in steering across models, with only 3% reduction in concept detection.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.06897 [cs.CL] |
| (or arXiv:2610.06897v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06897 arXiv-issued DOI via DataCite |
Submission history
From: Or Shafran [view email]
[v1]
Mon, 28 Sep 2026 17:57:46 UTC (835 KB)
来源:arXiv:cs.CL · arxiv.org