arXiv:cs.LG· Samy Houache (IMB, UB), Yann Traonmilin (IMB, UB), Jean-Fran\c{c}ois Aujol (UB, IMB)·· 4 小时前AI 评分35
分层误差归因实现快速鲁棒的混合精度训练后量化
Layerwise Error Attribution for Fast and Robust Mixed-Precision Post-Training Quantization
AI 导读
研究者提出分层误差归因方法,将量化误差拆分为传播误差与层内局部扰动,据此构建可分离评分,实现无需外部求解器的比特分配算法。在 DRUNet 去噪任务上,平均 4 比特预算下该方法的比特分配速度比基线快 28x 至 2,570x,校准数据被污染时 PSNR 提升最高达 7.5 dB。该框架直接应用于量化扩散模型也提升了 SOTA 表现。
正文
Abstract:Mixed-precision post-training quantization is a network compression method that assigns bits layer by layer, under a global memory budget using a small calibration set. The main difficulties are to overcome the combinatorial nature of the allocation problem and to manage the sensitivity to small, potentially corrupted databases. Hence, an efficient allocation method should be fast to compute and preserve model quality when calibration data are corrupted. To design such a method, we derive a layerwise probabilistic analysis of the quantization error that separates propagated error from the local perturbation introduced at a given layer. We use this local term to build a separable score for a simple allocation algorithm, that requires no external solver. The probabilistic nature of our approach brings robustness to corrupted data. On denoising tasks with DRUNet, with an average budget of 4 bits per weight, our method matches or improves state-of-the-art mixed-precision baselines under clean calibration, and is more robust to corrupted calibration, with PSNR gains of up to 7.5 dB under the tested corruptions. Experiments show bit-allocation speed-ups from 28x to 2,570x over the studied baselines. For quantized diffusion models, our experiments show that a direct application of our framework also improves the state-of-the-art.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09877 [cs.LG] |
| (or arXiv:2610.09877v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09877 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Samy Houache [view email] [via CCSD proxy]
[v1]
Wed, 7 Oct 2026 11:37:14 UTC (6,524 KB)
来源:arXiv:cs.LG · arxiv.org