跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Xinyu Wang, Sicheng Lyu, Xiao-Wen Chang·· 5 小时前AI 评分43

JARQ:面向量化的联合交替精修方法

JARQ: Joint Alternating Refinement for Quantization

AI 导读

JARQ 是一种即插即用的量化精修方法,从任意分组量化器出发,交替进行全部分组尺度的联合最小二乘拟合与有界 Babai 提议,可在不改变位宽、分组、零点与推理成本的前提下降低困惑度。

正文

View PDF HTML (experimental)

Abstract:Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that this leaves accuracy on the table: the best grid depends on the codes, input correlations couple the errors of different groups, and useful code changes often involve many codes at once. We propose JARQ , a plug-in refinement that starts from any group-wise quantizer and alternates a joint least-squares fit of all group scales with bounded Babai proposals that move many codes of a group together on the current grid. The problem is a bilinear box-constrained mixed-integer least-squares problem; the solver is backpropagation-free, does not increase the layer-wise objective under exact scale solves, and keeps the host's bit width, groups, zero points, and inference cost. Across Llama-2, Llama-3, and Qwen models with RTN, GPTQ, OmniQuant, and AWQ hosts, JARQ lowers perplexity in 90 of 96 comparisons, cuts three-bit RTN perplexity by up to 36%, raises mean multiple-choice accuracy in 23 of 24 configurations, and improves QEP, QuaRot, and OJBKQ outputs, at under a minute per 7B block.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.38599 [cs.LG]
  (or arXiv:2609.38599v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38599

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sicheng Lyu [view email]
[v1] Tue, 29 Sep 2026 22:00:00 UTC (680 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org