arXiv:cs.LG· Ziqun Bao, Xinyu Zhang, Yuchen Shao, Chengcheng Wan·· 4 小时前AI 评分55
TRACE:从 LLM checkpoint 更新中检测有害 SFT 痕迹的权重审计方法
Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates
AI 导读
arXiv 论文(arXiv:2610.07518)发现有害合规 SFT 会在 checkpoint 更新空间中留下连续、依赖目标的痕迹,坐标 s_H 在四个 7-8B 模型上与有害目标组成呈 0.986-0.992 的 Spearman 相关,且在更大模型规模上保持。
正文
Abstract:Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in checkpoint updates? We find that harmful-compliance SFT induces a continuous, objective-dependent ordering in checkpoint-update space. Using a reference geometry defined by pure harmful-compliance, safety-targeted, and benign-utility SFT, we find that a checkpoint-level coordinate s_H tracks controlled harmful-objective composition with Spearman correlations of 0.986-0.992 across four 7-8B backbones, with the same ordering persisting at larger model scales. Matched compliance-versus-refusal controls show that this checkpoint trace reflects the SFT objective rather than harmful-input exposure, while additional controls rule out simple explanations based on harmful-example count or generic training intensity. Building on this structure, we introduce TRACE, a weights-only auditing method that localizes an unknown checkpoint update relative to frozen harmful and non-harmful reference prototypes and converts this geometry into a continuous harmful-objective score. TRACE requires neither model queries nor access to the unknown SFT data, and can be evaluated directly from checkpoint updates. Across distribution shifts, unseen data, different SFT configurations, partial checkpoint access, and LoRA/full-parameter fine-tuning, the trace remains stable and is positively associated with independently measured attack success rates. TRACE remains informative even at low harmful-objective proportions, providing a complementary auditing signal when behavioral evaluation is unavailable or incomplete. Code is available at this https URL.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR) |
| Cite as: | arXiv:2610.07518 [cs.LG] |
| (or arXiv:2610.07518v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07518 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ziqun Bao [view email]
[v1]
Mon, 5 Oct 2026 23:28:45 UTC (936 KB)
来源:arXiv:cs.LG · arxiv.org