arXiv:cs.LG· Alexandra Volkova, Matin Ansaripour, Erik Schultheis, Christoph H. Lampert, Dan Alistarh·· 4 小时前AI 评分38
Q-PACE:面向量化感知训练的动态精度分配方法
Q-PACE: Dynamic Precision Allocation for Quantization-Aware Training
AI 导读
Q-PACE 是一种量化感知训练(QAT)的动态精度分配方法,通过二阶敏感度模型将损失增量建模为量化噪声 MSE 与逐层曲率系数加权之和。在最多 4B 参数的 LLM 预训练与监督微调实验中,Q-PACE 持续优于现有混合精度训练方案,并在显著更低的总内存预算下达到相当损失。研究还发现量化敏感度在深度和层类型上高度可预测,且训练中足够稳定,支持低频低成本的重校准。
正文
Abstract:Quantization-aware training (QAT) leverages lower-precision arithmetic to reduce the cost of LLM deployment, but aggressive quantization degrades final model performance. A common remedy is mixed-precision training, in which high precision is assigned to some of the layers to maintain performance while keeping the cost constrained. This approach then requires precision assignments for model layers during training. We provide a new approach, called Q-PACE, consisting of a second-order sensitivity model that predicts the loss increase as a sum of quantization noise MSE weighted by per-layer curvature coefficients. During training, we periodically re-compute these coefficients using perturbations across layers, and re-assign precision. Pretraining and supervised fine-tuning experiments on LLMs of up to 4B parameters show that Q-PACE consistently improves over existing mixed-precision training recipes, and achieves comparable loss at substantially lower total memory budgets. We further find that quantization sensitivity is highly predictable by depth and layer type, and its stability during training allows for infrequent, cheap recalibration.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09183 [cs.LG] |
| (or arXiv:2610.09183v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09183 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Alexandra Volkova [view email]
[v1]
Tue, 6 Oct 2026 22:34:48 UTC (1,298 KB)
来源:arXiv:cs.LG · arxiv.org