arXiv:cs.LG· Ashmitha R, J\"org Frochte·· 4 小时前AI 评分45
基于 Armijo 回溯的方向曲率:一种低成本锐度探测与 Adam 免校准学习率保护机制
Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam
AI 导读
研究者提出用 Armijo 回溯线搜索中接受的步长 α 来估计方向曲率,无需运行 Lanczos 或 Hessian-向量乘积。在 CIFAR-10、Fashion-MNIST 和 Imagenette 上,log α 与最大 Hessian 特征值 λ₁ 的 Pearson 相关性达 -0.91 至 -0.95。
正文
Abstract:Local sharpness, defined by the largest Hessian eigenvalue $\lambda_1$, sets the maximum stable gradient update size, but its computation would usually require running Lanczos or Hessian-vector products. However, we notice that even a single Armijo backtracking line search already contains this information with just a few forward passes, as the accepted step $\alpha$ determines the directional curvature along the search direction up to the multiplicative band set by the backtracking factor. The correlation between $\log\alpha$ and $\log\lambda_1$ on CIFAR-10, Fashion-MNIST and Imagenette reaches $-0.91$ to $-0.95$ in Pearson correlation, and even after removing the trend per run the correlation remains at $-0.60$ to $-0.70$. This allows for a cheap online Edge-of-Stability estimate of the slow sharpness component. The employed probing mechanism searches along Adam's first-step update direction, at initialisation and nine times over the course of the first 50 optimiser steps. The learning-rate cap is set as twice the smallest observed step size, and in the studied learning-rate ranges ($10^{-3}$ to $3.0$) and GPT-2 pretraining experiments all capped runs avoid divergence; in the general architecture analysis, one MLP architecture is still sensitive to the initial batch order. This probing protocol incurs an approximately one percent overhead, and using a non-binding cap means that the optimiser's state and first update are bit-identical. There is no fine-tuning of any of the protocol parameters to any specific architecture; this is the sense in which this is a calibration-free safeguard. It is meant as a way of avoiding divergence, not achieving accuracy. At GPT-2 scale, multiple measurements also illustrate why an initialisation-only cap is insufficient: directional curvature rises substantially in the first five optimiser steps, which motivates the short probationary window.
| Comments: | 38 pages, 7 figures. Published in Transactions on Machine Learning Research (10/2026): this https URL. Code: this https URL |
| Subjects: | Machine Learning (cs.LG) |
| MSC classes: | 68T07, 90C26, 65K05 |
| ACM classes: | I.2.6; G.1.6 |
| Cite as: | arXiv:2607.03998 [cs.LG] |
| (or arXiv:2607.03998v5 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2607.03998 arXiv-issued DOI via DataCite |
|
| Journal reference: | Transactions on Machine Learning Research (ISSN 2835-8856), October 2026 |
Submission history
From: Jörg Frochte [view email]
[v1]
Sat, 4 Jul 2026 19:49:01 UTC (732 KB)
[v2]
Sat, 11 Jul 2026 20:23:10 UTC (733 KB)
[v3]
Wed, 15 Jul 2026 20:08:42 UTC (733 KB)
[v4]
Sun, 16 Aug 2026 08:02:20 UTC (808 KB)
[v5]
Wed, 7 Oct 2026 09:44:06 UTC (810 KB)
来源:arXiv:cs.LG · arxiv.org