跳到正文
arXiv:cs.LG· Ashmitha R, J\"org Frochte·· 4 小时前AI 评分45

基于 Armijo 回溯的方向曲率:一种低成本锐度探测与 Adam 免校准学习率保护机制

Directional Curvature from Armijo Backtracking: A Low-Cost Sharpness Probe and a Calibration-Free Learning-Rate Safeguard for Adam

AI 导读

研究者提出用 Armijo 回溯线搜索中接受的步长 α 来估计方向曲率,无需运行 Lanczos 或 Hessian-向量乘积。在 CIFAR-10、Fashion-MNIST 和 Imagenette 上,log α 与最大 Hessian 特征值 λ₁ 的 Pearson 相关性达 -0.91 至 -0.95。

正文

View PDF

Abstract:Local sharpness, defined by the largest Hessian eigenvalue $\lambda_1$, sets the maximum stable gradient update size, but its computation would usually require running Lanczos or Hessian-vector products. However, we notice that even a single Armijo backtracking line search already contains this information with just a few forward passes, as the accepted step $\alpha$ determines the directional curvature along the search direction up to the multiplicative band set by the backtracking factor. The correlation between $\log\alpha$ and $\log\lambda_1$ on CIFAR-10, Fashion-MNIST and Imagenette reaches $-0.91$ to $-0.95$ in Pearson correlation, and even after removing the trend per run the correlation remains at $-0.60$ to $-0.70$. This allows for a cheap online Edge-of-Stability estimate of the slow sharpness component. The employed probing mechanism searches along Adam's first-step update direction, at initialisation and nine times over the course of the first 50 optimiser steps. The learning-rate cap is set as twice the smallest observed step size, and in the studied learning-rate ranges ($10^{-3}$ to $3.0$) and GPT-2 pretraining experiments all capped runs avoid divergence; in the general architecture analysis, one MLP architecture is still sensitive to the initial batch order. This probing protocol incurs an approximately one percent overhead, and using a non-binding cap means that the optimiser's state and first update are bit-identical. There is no fine-tuning of any of the protocol parameters to any specific architecture; this is the sense in which this is a calibration-free safeguard. It is meant as a way of avoiding divergence, not achieving accuracy. At GPT-2 scale, multiple measurements also illustrate why an initialisation-only cap is insufficient: directional curvature rises substantially in the first five optimiser steps, which motivates the short probationary window.
Comments: 38 pages, 7 figures. Published in Transactions on Machine Learning Research (10/2026): this https URL. Code: this https URL
Subjects: Machine Learning (cs.LG)
MSC classes: 68T07, 90C26, 65K05
ACM classes: I.2.6; G.1.6
Cite as: arXiv:2607.03998 [cs.LG]
  (or arXiv:2607.03998v5 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2607.03998

arXiv-issued DOI via DataCite

Journal reference: Transactions on Machine Learning Research (ISSN 2835-8856), October 2026

Submission history

From: Jörg Frochte [view email]
[v1] Sat, 4 Jul 2026 19:49:01 UTC (732 KB)
[v2] Sat, 11 Jul 2026 20:23:10 UTC (733 KB)
[v3] Wed, 15 Jul 2026 20:08:42 UTC (733 KB)
[v4] Sun, 16 Aug 2026 08:02:20 UTC (808 KB)
[v5] Wed, 7 Oct 2026 09:44:06 UTC (810 KB)

来源:arXiv:cs.LG · arxiv.org