跳到正文
arXiv:cs.LG· Domenic Rosati, Alessa Carbo, Ali Dadsetan, Hong Huang, Matthew Young, Subhabrata Majumdar, Frank Rudzicz, Hassan Sajjad·· 4 小时前AI 评分40

移除信息内容不能证明开放权重模型的抗篡改能力

Removing Information Content Does Not Certify Tamper Resistance in Open-Weight Models

AI 导读

开放权重模型移除有害信息后,其发布时的互信息不足以普遍证明模型能抵抗微调攻击。研究发现,保函数的重参数化可在互信息不变的情况下改变梯度下降几何结构,训练顺序能在固定权重-数据互信息下改变恢复时间,而表征层面的精确独立甚至可保留全部参数雅可比矩阵;论文给出的一种显式构造使两项互信息量均为零,模型仅需一步梯度即可恢复。

正文

View PDF HTML (experimental)

Abstract:Does removing harmful information make open-weight models resistant to fine-tuning attacks? We show that mutual information at release alone cannot universally certify slow recovery. Function-preserving reparameterizations leave information unchanged while altering gradient-descent geometry, so an invariant certificate is bounded by the fastest reachable parameterization. We apply this principle to weight--data mutual information under training-data filtering and label--representation mutual information under capability removal. Training order can change recovery time at fixed weight--data information, while exact representation-level independence can preserve the entire parameter Jacobian. An explicit construction has both information quantities equal to zero and recovers in one gradient step. Controlled experiments illustrate order-dependent recovery and parameterization-dependent attack speed. These results identify the missing requirement for certification: constraints on attack dynamics beyond mutual information at release.
Comments: Under submission AISTATS 2026
Subjects: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
Cite as: arXiv:2610.09004 [cs.LG]
  (or arXiv:2610.09004v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09004

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Domenic Rosati [view email]
[v1] Tue, 6 Oct 2026 18:59:57 UTC (118 KB)

来源:arXiv:cs.LG · arxiv.org