跳到正文
arXiv:cs.LG· Minyoung Choi, Dalta Imam Maulana, Wanyeong Jung·· 4 小时前AI 评分40

轻量通用学习型优化器:通过梯度历史重组实现

Lightweight and Versatile Learned Optimization by Recombination of Gradient History

AI 导读

一种轻量通用学习型优化器通过动态重组梯度历史(以不重叠时间段的均值表示)来降低预测空间。它用37k参数网络、0.87 GPU小时训练,零样本泛化到未见任务,在BERT-Tiny和GPT-Tiny上将验证损失分别降低9.1%和0.4%,在Vision Transformer上比Adam提升3.5%p测试准确率,九个图模型平均提升2.7%p,FLOPs开销低至0.3%。

正文

View PDF HTML (experimental)

Abstract:This paper presents a lightweight and versatile learned optimizer that dynamically recombines gradient history, represented as averages over disjoint time spans. The optimizer reduces the prediction space to one scalar coefficient per gradient average, shared by multiple parameters. Progressively averaging older gradients minimizes memory cost of long history, while keeping their contributions independently accessible. A 37k-parameter network trained in 0.87 GPU-hours generalizes zero-shot to unseen tasks, lowering validation loss by 9.1% and 0.4% on BERT-Tiny and GPT-Tiny, and improving test accuracy over Adam by 3.5 %p on a Vision Transformer and by 2.7 %p on average across nine graph models, with FLOPs overhead as low as 0.3%.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.09604 [cs.LG]
  (or arXiv:2610.09604v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09604

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Minyoung Choi [view email]
[v1] Wed, 7 Oct 2026 07:50:55 UTC (3,339 KB)

来源:arXiv:cs.LG · arxiv.org