跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Vicente Opazo, Jose Calatayud-Mateu, Cristobal Rojas, Cristian Buc Calderon·· 5 小时前AI 评分40

平坦度代理不等于函数:鲁棒性证书与训练干预的失效分析

When a Flatness Proxy Is Not a Function: Robustness Certificates and Training Interventions

AI 导读

研究发现,最后一层相对平坦度代理虽能给出有效的曲率上界,却不足以支撑鲁棒性证书或训练干预。在有限全局经验风险最小值处,保留的证书表达式将损失增幅低估超过 210 倍;常见行 softmax 平移会保持预测与收缩性,却使该代理无界。45 组配对单步测试中,放大平移可区分原始正则化与商正则化预测器,CIFAR-10 长时程实验显示泛化能力被大幅可逆抑制。

正文

View PDF HTML (experimental)

Abstract:A valid curvature upper bound need not justify either a robustness certificate or an intervention on an intrinsic predictor property. We demonstrate this distinction for a last-layer relative-flatness proxy used in both settings. First, empirical-risk stationarity does not eliminate pointwise first-order loss terms: at a finite global empirical-risk minimum, the retained certificate expression underestimates a loss increase by over $210\times$. We derive a globally valid, gauge-invariant feature-space repair. Second, common-row softmax shifts preserve predictions and the exact contraction while making the proxy unbounded. Even standard reference-class choices double it on average relative to the centered representation. For a single fixed-feature example with at least three classes, scalar retuning generically cannot align the induced probability updates. Row centering gives the orbit-minimized bound and restores value and full-model gradient invariance under this symmetry. Across 45 paired one-step tests on algorithmic and image models, amplified shifts separate raw-regularized predictors while quotient-regularized predictors remain aligned. Long-horizon CIFAR-10 experiments show substantial, reversible suppression of generalization, while evidence for selective delay after memorization is less consistent. Together, these results show that validity as a curvature upper bound does not by itself justify either inversion into a robustness certificate or differentiation into an intrinsic training intervention.
Comments: 10 pages, 3 figures, 1 table
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.38540 [cs.LG]
  (or arXiv:2609.38540v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38540

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Vicente Opazo [view email]
[v1] Tue, 29 Sep 2026 20:57:34 UTC (45 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org