跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Jonathan Mei, Sang Hyub Kim, Oliver Knitter, Chi Chen, Martin Roetteler·· 5 小时前AI 评分39

ShamAN-Q:面向亚 1-bit LLM 权重的 Shampoo 增强 NanoQuant 量化方法

ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM Weights

AI 导读

ShamAN-Q 是一种亚 1-bit 训练后量化方法,通过用可处理的稠密曲率度量替换 NanoQuant 的对角重构几何,将 Shampoo 优化器的通用范式引入量化。

正文

View PDF HTML (experimental)

Abstract:We introduce ShamAN-Q, a sub-1-bit post-training quantization method that extends NanoQuant by replacing each its diagonal reconstruction geometry with a tractable dense curvature metric, using a general paradigm popularized by the Shampoo optimizer. For each linear weight, ShamAN-Q fits a Kronecker product to the empirical Fisher information matrix of a small calibration set by Kullback--Leibler minimization, forming a Mahalanobis reconstruction loss from the result. The continuous ADMM updates from NanoQuant become solutions to Sylvester equations, while its discrete projection and deployment format remain unchanged. Because the curvature is local to a given set of weights, ShamAN-Q re-measures the input curvature statistic for each layer immediately before layer factorization, periodically refreshing all statistics on the partially quantized model. ShamAN-Q also redistributes the uniform rank from NanoQuant across layers at the same total number of bits. On Qwen3-Base, ShamAN-Q lowers WikiText-2 perplexity at $\approx$1 bpw from 27.56 to 22.96 (0.6B), 19.21 to 16.72 (1.7B), and 14.29 to 13.80 (4B) while matching or improving zero-shot accuracy on the Eleuther LM Evaluation Harness. On 0.6B, ShamAN-Q at $\approx$0.8 bpw matches the published perplexity of NanoQuant at $\approx$1.0 bpw.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as: arXiv:2609.38521 [cs.LG]
  (or arXiv:2609.38521v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38521

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jonathan Mei [view email]
[v1] Tue, 29 Sep 2026 20:39:47 UTC (343 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org