跳到正文
arXiv:cs.LG· Md Sazid Uddin, Md. Khairul Alam Mazumder, M. F. Mridha·· 3 小时前AI 评分44

认证式机制编辑:技能移除与保留的行为保证

Certified Mechanistic Edits: Behavioral Guarantees for Skill Removal and Preservation

AI 导读

研究者提出"认证式机制编辑",可证明性地保证在连续嵌入空间区域内禁用某回路即移除一项技能、同时保留另一项技能,从玩具 ReLU 网络一路验证到标准 softmax + LayerNorm Transformer,并借助可靠的边界传播将可处理的输入扰动维度提升约 9 倍。

正文

View PDF HTML (experimental)

Abstract:Mechanistic edits (ablations, weight edits, activation steering) are the standard tools for unlearning a harmful capability from a neural network while preserving useful ones. Current approaches validate their effects only by testing, which can never cover an entire continuous region of inputs. Prior work at the interpretability-verification boundary certifies descriptions of a model: what a circuit computes, or whether it faithfully explains the whole. We instead certify the behavioral effect of an edit: that disabling a circuit removes one skill and provably preserves another, for every input in a region; a feature non-interference guarantee in the information-flow-security sense. We demonstrate such certified edits from toy ReLU networks up to a standard softmax + LayerNorm transformer, proving removal and preservation over continuous embedding-space regions and reaching roughly 9x the input-perturbation dimension an exact solver can handle by switching to sound bound propagation. Furthermore, we prove that no finite deterministic black-box test can certify removal, exhibiting an edit that passes exhaustive testing yet provably fails on a survivor pocket that can be made arbitrarily small. Guarantees hold on small, standard-architecture networks and, like any removal claim, presuppose that the target skill admits a decidable specification, a property which real-world harms may not have.
Comments: 12 pages, 6 figures, 4 tables
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
MSC classes: 68T07, 68Q60
ACM classes: I.2.6; D.2.4; F.3.1
Cite as: arXiv:2610.03502 [cs.LG]
  (or arXiv:2610.03502v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03502

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Md Sazid Uddin [view email]
[v1] Fri, 2 Oct 2026 16:01:23 UTC (129 KB)

来源:arXiv:cs.LG · arxiv.org