跳到正文
arXiv:cs.CL· Muhammad Zeeshan Karamat, Christiana Chamon Garcia·· 3 小时前AI 评分43

LLaMA-2-7B-Chat 的端侧安全有多脆弱?定位稀疏故障分析的安全关键参数

How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis

AI 导读

研究考察 LLaMA-2-7B-Chat 的安全敏感行为是否集中在稀疏参数子集中。低秩安全子空间分析与参数级安全—效用重要性过滤均显示安全敏感性高度非均匀,MLP 的 down_proj 是最突出的安全敏感组件,o_proj 贡献较小。

正文

View PDF HTML (experimental)

Abstract:As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern. We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis. We study two complementary localization methods: low-rank safety-associated subspace analysis and parameter-level safety--utility importance filtering. Both approaches reveal highly non-uniform safety sensitivity across the network, with the MLP down_proj consistently emerging as a prominent safety-sensitive component and o_proj providing a smaller contribution. Using parameter-level localization, modifying only 0.19% of model weights in down_proj yields 53% Basic ASR and 56% GCG ASR, while tinyBenchmarks accuracy remains at 51.6% compared with a 52.2% unmodified baseline. These results motivate targeted fault analysis and selective integrity protection for language models deployed in resource-constrained, on-device, and agentic settings.
Comments: Accepted at NeurIPS 2026 Workshop
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Software Engineering (cs.SE)
Cite as: arXiv:2610.09000 [cs.AI]
  (or arXiv:2610.09000v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.09000

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Muhammad Zeeshan Karamat [view email]
[v1] Tue, 6 Oct 2026 18:56:14 UTC (17 KB)

来源:arXiv:cs.CL · arxiv.org