arXiv:cs.CL· Kai Yi, Tarek Elgamal, Sruthikesh Surineni, Vignesh Vivekraja, Soumyadeep Ghosh, Steven Li·· 3 小时前AI 评分43
CanonQ:面向 W2A4KV2 极端低比特 LLM 压缩的统一量化感知训练框架
Few Bits, One Law: Toward W2A4KV2
AI 导读
研究者提出 CanonQ,一个统一量化感知训练框架,通过将源规范化与任务自适应分离,把异构张量源映射到规范坐标,使冻结的高斯参考码本可跨层跨模型复用。在最强的 W2A4KV2 联合压缩下,CanonQ-Omni 在 LLaMA3-1B/3B/8B 上实现最高 14.28 倍的 WikiText-2 困惑度降低和最高 57.9% 的零样本平均准确率提升。
正文
Abstract:Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantization errors interact throughout the network. We introduce CanonQ, a unified quantization-aware training framework that addresses these challenges by separating source canonicalization from task-aware adaptation. Fixed rotations and energy normalization map heterogeneous tensor sources to canonical coordinates, enabling frozen Gaussian-reference codebooks to be reused across layers and models. Joint training then adapts the network to the coupled errors of weight, activation, and cache quantization within a common scalar/vector interface. We bound frozen-codebook transfer error and local task loss, and derive an exact normalization-aware straight-through Jacobian that links quantization distortion to gradient bias. The strongest gains arise under joint W2A4KV2 compression: across LLaMA3-1B/3B/8B, CanonQ-Omni achieves up to 14.28x lower WikiText-2 perplexity and up to 57.9% higher mean zero-shot accuracy than prior state-of-the-art and representative quantization baselines. The benefits extend to Qwen3-1.7B, code generation, and mathematical reasoning: on instruction-tuned MobileLLM-Pro-1B at W2A16KV16, CanonQ achieves relative improvements of 41.7% in HumanEval pass@1 and 39.1% in GSM8K exact match over the strongest evaluated quantization baseline.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09202 [cs.AI] |
| (or arXiv:2610.09202v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09202 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kai Yi [view email]
[v1]
Tue, 6 Oct 2026 22:55:36 UTC (564 KB)
来源:arXiv:cs.CL · arxiv.org