arXiv:cs.LG· Vincent Counathe, Ben Athiwaratkun, Christopher De Sa, Tianyi Zhang·· 7 小时前AI 评分40
QUASAR:用损失感知重建降低量化感知训练损失下界
QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction
AI 导读
QUASAR 是一种将轻量级损失感知重建引入训练循环的量化感知训练(QAT)方法,通过在小范围 clipping 区间搜索并以显著度加权最小二乘拟合反量化参数来重建潜在权重。
正文
Abstract:As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model quality. However, QAT has a structural mismatch: gradient updates are applied to latent full-precision weights, while the loss and gradients are computed on lossy reconstructions of those weights. This mismatch can lead to suboptimal training trajectories and a higher loss floor. Second-order PTQ methods address a similar problem by minimizing loss-aware reconstruction error, but applying such expensive reconstruction repeatedly during QAT as the weights evolve is impractical. We introduce QUASAR, a QAT method that brings lightweight, loss-aware reconstruction into the training loop. At each training step, QUASAR reconstructs the latent weights by searching over a small set of clipping ranges and fitting dequantization parameters through saliency-weighted least squares. We use an exponential moving average of squared gradients as the per-parameter saliency signal. Our theoretical analysis shows that optimizing QUASAR's reconstruction objective tightens both the convergence and final-loss bounds of QAT. We evaluate QUASAR across four model families and across INT4, INT3, INT2, and NVFP4 quantization formats. QUASAR consistently achieves lower training and evaluation loss than competitive QAT methods and outperforms QAT and PTQ baselines on downstream benchmarks. At INT2, QUASAR improves average accuracy over the best QAT baseline by 13.3 points with quantization-aware distillation and by 10.9 points with QAT on mathematical reasoning data. Notably, after distillation on only about 600M tokens, QUASAR's INT4 Gemma-4 E4B checkpoint outperforms the corresponding QAT checkpoint released by Google, with 66% lower KL divergence and 1.8 points higher average accuracy.
| Comments: | 44 pages |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL); Machine Learning (stat.ML) |
| Cite as: | arXiv:2608.13966 [cs.LG] |
| (or arXiv:2608.13966v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2608.13966 arXiv-issued DOI via DataCite |
Submission history
From: Vincent Counathe [view email]
[v1]
Fri, 14 Aug 2026 05:29:58 UTC (1,348 KB)
[v2]
Mon, 5 Oct 2026 21:37:32 UTC (2,807 KB)
来源:arXiv:cs.LG · arxiv.org