arXiv:cs.LG· Md Sazid Uddin, Md. Khairul Alam Mazumder, M. F. Mridha·· 3 小时前AI 评分44
认证式机制编辑:技能移除与保留的行为保证
Certified Mechanistic Edits: Behavioral Guarantees for Skill Removal and Preservation
AI 导读
研究者提出"认证式机制编辑",可证明性地保证在连续嵌入空间区域内禁用某回路即移除一项技能、同时保留另一项技能,从玩具 ReLU 网络一路验证到标准 softmax + LayerNorm Transformer,并借助可靠的边界传播将可处理的输入扰动维度提升约 9 倍。
正文
Abstract:Mechanistic edits (ablations, weight edits, activation steering) are the standard tools for unlearning a harmful capability from a neural network while preserving useful ones. Current approaches validate their effects only by testing, which can never cover an entire continuous region of inputs. Prior work at the interpretability-verification boundary certifies descriptions of a model: what a circuit computes, or whether it faithfully explains the whole. We instead certify the behavioral effect of an edit: that disabling a circuit removes one skill and provably preserves another, for every input in a region; a feature non-interference guarantee in the information-flow-security sense. We demonstrate such certified edits from toy ReLU networks up to a standard softmax + LayerNorm transformer, proving removal and preservation over continuous embedding-space regions and reaching roughly 9x the input-perturbation dimension an exact solver can handle by switching to sound bound propagation. Furthermore, we prove that no finite deterministic black-box test can certify removal, exhibiting an edit that passes exhaustive testing yet provably fails on a survivor pocket that can be made arbitrarily small. Guarantees hold on small, standard-architecture networks and, like any removal claim, presuppose that the target skill admits a decidable specification, a property which real-world harms may not have.
| Comments: | 12 pages, 6 figures, 4 tables |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR) |
| MSC classes: | 68T07, 68Q60 |
| ACM classes: | I.2.6; D.2.4; F.3.1 |
| Cite as: | arXiv:2610.03502 [cs.LG] |
| (or arXiv:2610.03502v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03502 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Md Sazid Uddin [view email]
[v1]
Fri, 2 Oct 2026 16:01:23 UTC (129 KB)
来源:arXiv:cs.LG · arxiv.org