arXiv:cs.LG· Yan Scholten, Rachel Lawrence, James Hensman, Stephan G\"unnemann, Alicia Curth, Riccardo Grazzi·· 4 小时前AI 评分34
激活去噪:并行与串行 LLM 量化的鲁棒性视角对比
Activation Denoising: A Robustness View on Parallel vs Sequential LLM Quantization
AI 导读
研究者提出"激活去噪"方法,让 LLM 量化在完全并行的同时恢复串行量化的大部分收益,耗时仅为后者的一小部分。该方法不再逐层重新校准,而是将上游量化误差建模为噪声,通过预处理与度量加权舍入实现鲁棒正则化,形成随深度累积的平滑惩罚。这种鲁棒正则化与量化中常用的正交旋转互补,效果可叠加。
正文
Abstract:Post-training quantization is a powerful tool for compressing large language models. The most scalable methods quantize every layer in parallel, but quantization errors then compound through the residual stream, as no layer corrects for the errors of the layers before it. Sequential quantization accounts for this error compounding by re-calibrating each layer on the already-quantized outputs of its predecessors, yielding stronger results but at the cost of a serial schedule that becomes a bottleneck at scale. As a solution, we propose parallel quantization with activation denoising, which recovers much of the sequential benefit while keeping quantization fully parallel. Rather than re-calibrating layer-by-layer, we take a robustness perspective and model the upstream error as noise, regularizing to be robust to it through a preprocessing step followed by metric-weighted rounding. Applied at every layer, this regularization forms a depth-compounding smoothness penalty that dampens how strongly quantization errors amplify through the model. Unlike orthogonal rotations commonly used in quantization, which must preserve the model's function, we multiply the weights by a more general linear transformation. We find that the two are complementary and their effects compound. Empirically, our robustness regularization recovers a significant part of sequential quantization's benefit in a single parallel pass, at a fraction of its time. Overall, by treating compounding quantization errors as a robustness problem, we offer a principled foundation for more efficient and accurate LLM quantization at scale.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07522 [cs.LG] |
| (or arXiv:2610.07522v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07522 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yan Scholten [view email]
[v1]
Mon, 5 Oct 2026 23:41:48 UTC (610 KB)
来源:arXiv:cs.LG · arxiv.org