arXiv:cs.LG· Abdul Kadir (University of Oldenburg, Oldenburg, Germany, German Research Center for Artificial Intelligence), Md Mohasin Hossain (German Research Center for Artificial Intelligence, Saarland University, Saarbrucken, Germany), Daniel Sonntag (University of Oldenburg, Oldenburg, Germany, German Research Center for Artificial Intelligence)·· 4 小时前AI 评分35
语言模型中负责网络信息检索的注意力头与神经元定位研究
Finding the Heads and the Neurons Responsible for Network Information Retrieval in Language Models
AI 导读
研究因果验证了语言模型识别网络基础设施信息(主机名与 IP 地址配对)的责任是否集中于特定注意力头及头内神经元,覆盖三个架构家族的五个模型。在每个模型中,少数注意力头(128 至 1152 个候选中的 1 至 9 个)即可支撑检测器达到 99.5%–100% 的留出准确率。
正文
Authors:Abdul Kadir (1 and 2), Md Mohasin Hossain (2 and 3), Daniel Sonntag (1 and 2) ((1) University of Oldenburg, Oldenburg, Germany, (2) German Research Center for Artificial Intelligence (DFKI), Germany, (3) Saarland University, Saarbrucken, Germany)
Abstract:We ask whether specific attention heads, and more finely specific neurons inside those heads, are responsible for recognizing that a language model's context contains network infrastructure information (a hostname paired with its IP address), and whether that responsibility can be validated causally rather than by correlation alone. At the head level the answer is yes, across five models spanning three architecture families: in every model, a small set of heads (1 to 9 out of 128 to 1152 candidates), found by causal ablation screening and tested for selectivity against matched negative and context-free controls, supports a detector with 99.5--100\% held-out accuracy. We then ask whether a head's responsibility concentrates into one neuron or stays spread across its dimensions; this is model-specific. In one model, the top head's signal concentrates into a single neuron, found independently by both a causal intervention and a correlational ranking, which agree exactly (AUC = 1.000, matching the full head). In another, the single clean head works as a whole (AUC = 1.000) but the best causally ranked neuron inside it does not (AUC = 0.665), so the responsibility there is spread across the head. The remaining three models fall in between. On an independent dataset collected by a different institution (reverse-DNS records rather than the discovery data), every model's full-head detector flags 100\% of positive records; the single-neuron versions transfer less reliably, and in one model score below chance. Causal head-finding for a specific network-information entity works across models and architectures; how far that finding can be pushed down to individual neurons varies, and needs to be checked for each model.
| Comments: | 13 pages, 2 figures, 12 tables |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.08200 [cs.LG] |
| (or arXiv:2610.08200v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08200 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Md Mohasin Hossain [view email]
[v1]
Tue, 6 Oct 2026 11:51:03 UTC (47 KB)
来源:arXiv:cs.LG · arxiv.org