arXiv:cs.CL· Muhammad Zeeshan Karamat, Christiana Chamon Garcia·· 3 小时前AI 评分43
LLaMA-2-7B-Chat 的端侧安全有多脆弱?定位稀疏故障分析的安全关键参数
How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis
AI 导读
研究考察 LLaMA-2-7B-Chat 的安全敏感行为是否集中在稀疏参数子集中。低秩安全子空间分析与参数级安全—效用重要性过滤均显示安全敏感性高度非均匀,MLP 的 down_proj 是最突出的安全敏感组件,o_proj 贡献较小。
正文
Abstract:As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern. We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis. We study two complementary localization methods: low-rank safety-associated subspace analysis and parameter-level safety--utility importance filtering. Both approaches reveal highly non-uniform safety sensitivity across the network, with the MLP down_proj consistently emerging as a prominent safety-sensitive component and o_proj providing a smaller contribution. Using parameter-level localization, modifying only 0.19% of model weights in down_proj yields 53% Basic ASR and 56% GCG ASR, while tinyBenchmarks accuracy remains at 51.6% compared with a 52.2% unmodified baseline. These results motivate targeted fault analysis and selective integrity protection for language models deployed in resource-constrained, on-device, and agentic settings.
| Comments: | Accepted at NeurIPS 2026 Workshop |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Software Engineering (cs.SE) |
| Cite as: | arXiv:2610.09000 [cs.AI] |
| (or arXiv:2610.09000v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09000 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Muhammad Zeeshan Karamat [view email]
[v1]
Tue, 6 Oct 2026 18:56:14 UTC (17 KB)
来源:arXiv:cs.CL · arxiv.org