arXiv:cs.AI· Tong Che, Yilong Li·· 6 小时前AI 评分45
LLMs 中的特征:模型是否真正使用了它?区分引导与机制
Does the Model Use the Feature? Separating Steering from Mechanism in LLMs
AI 导读
研究提出一种经验契约,通过在自然输入观测值上测试特征来区分"引导"与"机制"。安装测试衡量特征对行为的充分性,移除与下游救援测试衡量模型对该特征的使用程度。在 LLM 三类表征上,两种强度差异显著:已发表的未知实体潜变量虽能强烈引导知识弃答,但安装观测值仅传递自然已知-未知弃答对比的一小部分。
正文
Abstract:Internal features in LLMs are often interpreted as mechanisms when they track a concept and their manipulation changes a related behavior. Yet steering can push a feature far outside its natural range, where its effects need not reflect the model's own computation. We examine this inference and propose an empirical contract whose tests evaluate features at values observed on natural inputs. One test copies a feature's value from an input that shows a behavior into a matched input that does not (installation) or the reverse (removal); the other restores the feature after an upstream edit (downstream rescue). Installation measures how far the feature suffices for the behavior; removal and downstream rescue measure how much the model uses it. Applied to three kinds of representations, the two strengths separate sharply. The published unknown-entity latent strongly steers knowledge abstention, yet installing observed values from either published latent into matched prompts transfers only a small fraction of the natural known--unknown abstention contrast. Dense known--unknown directions show opposite asymmetries between installation and removal in Gemma and Llama, and how fully a released subject--verb agreement feature set reproduces and restores the behavior depends on how its values are written into the model. Tracking a concept and steering a behavior therefore do not by themselves show that the model uses a feature, and each conclusion holds only for the intervention tested.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07270 [cs.AI] |
| (or arXiv:2610.07270v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07270 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yilong Li [view email]
[v1]
Mon, 5 Oct 2026 19:11:42 UTC (122 KB)
来源:arXiv:cs.AI · arxiv.org