arXiv:cs.LG· Bishnu Bhusal, Minh Vu, Ben Southworth, Geigh Zollicoffer, Rohit Chadha, Manish Bhattarai·· 3 小时前AI 评分36
面向差分隐私 Muon 的内动量方法
Inner Momentum for Differentially Private Muon
AI 导读
针对差分隐私训练中逐样本梯度裁剪会扭曲 Muon 更新所依赖的奇异向量几何,作者提出 DP-Muon-IM:在裁剪前将每个样本的 Muon 梯度在当前模型与近期模型历史上做平均。
正文
Abstract:Differentially private training clips each per-example gradient before adding noise. This clipping is radial for each example, yet unequal clipping factors can distort the relative singular-vector geometry of their average. Muon is particularly exposed to this effect, since its update is an approximate polar factor UV^T that depends only on the singular vectors that clipping can shift. To curb this degradation, we propose averaging each sampled example's Muon gradient over the current model and a short history of recent models before clipping. The clipped batch matrix then separates into a common rescaling and a covariance residual R between sampled gradients and clipping values, with ||R||_F <= sigma_lambda sigma_G, bounding the clipping-induced distortion directly. We further show that a finite Newton-Schulz iteration preserves the polar factor of its input under these spectral conditions, confirming that our correction survives orthogonalization. In private GPT-2 fine-tuning on E2E and DART at epsilon in {1, 2, 4, 8}, DP-Muon-IM improves BLEU and ROUGE-L over DP-Muon in every seed-matched comparison, and non-private diagnostics show 2-4% lower pre-noise polar error.
| Comments: | LA-UR Number: LA-UR-26-28799 |
| Subjects: | Machine Learning (cs.LG); Cryptography and Security (cs.CR) |
| Cite as: | arXiv:2610.02738 [cs.LG] |
| (or arXiv:2610.02738v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02738 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Manish Bhattarai [view email]
[v1]
Fri, 2 Oct 2026 03:10:53 UTC (281 KB)
来源:arXiv:cs.LG · arxiv.org