跳到正文
arXiv:cs.LG· Hanyang Li, Shao Tang, Daniel Thomas Braithwaite, Gregory Dexter, Leonardo Neves, Aman Gupta, Hiroto Udagawa, Abhishek Shivanna, Daniel Silva, Rohan Ramanath·· 3 小时前

在预条件子空间做舍入:4-bit AdamW 优化器状态量化重新设计

Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization

AI 导读

研究者提出 ZIP-SR 与 ZE-EDEN 两种 4-bit AdamW 优化器状态量化方案,从"舍入空间"视角重新设计量化过程。两者第一动量均用 4-bit NF4,并在训练最后 10% 对 LM-head 第一动量做定向随机舍入。

正文

View PDF HTML (experimental)

Abstract:Quantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates. We redesign 4-bit optimizer-state quantization for AdamW from the perspective of \emph{rounding space}: the coordinate in which a quantizer chooses between adjacent reconstruction levels. For the second moment, a local analysis of the quantization cell adjacent to zero shows that small mean state error need not imply small mean preconditioner error at the next step. A one-dimensional quadratic construction further shows qualitatively different optimization dynamics under state-space and preconditioner-space rounding. These results motivate Zero-Inclusive Preconditioner-space Stochastic Rounding (\textbf{ZIP-SR}), which retains zero in the second-moment codebook and computes stochastic-rounding probabilities in preconditioner space. As a complementary route, Zero-Excluding EDEN calibration (\textbf{ZE-EDEN}) uses a zero-excluding second-moment codebook and rescales the quantized second-moment block to mitigate the preconditioner distortion caused by the positive quantization floor. Both configurations use 4-bit NormalFloat (NF4) for the first moment, with targeted stochastic rounding of the LM-head first moment during the final 10\% of training. Across GPT- and Llama-style pretraining experiments ranging from \textbf{130M} to \textbf{2.7B} parameters, both methods reduce TorchAO 4-bit AdamW's mean validation-loss gap to 32-bit AdamW at every evaluated model size, with the largest reported gap reduction reaching \textbf{70\%}. In full-parameter supervised fine-tuning, both recipes achieve lower validation loss than TorchAO while remaining close to 32-bit AdamW on downstream tasks.
Comments: 23 pages
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.12444 [cs.LG]
  (or arXiv:2610.12444v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.12444

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Aman Gupta [view email]
[v1] Thu, 8 Oct 2026 17:58:09 UTC (294 KB)

来源:arXiv:cs.LG · arxiv.org