arXiv:cs.LG· Guanhua Ding, Zi Wang, Ruichao Li, Jack Liu·· 4 小时前AI 评分42
CurveTQ:无需旋转的 LLM 权重 Trellis 量化,基于曲率加权搜索
CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search
AI 导读
CurveTQ 提出一种无需随机正交旋转的 LLM 权重 Trellis 量化方法,将 Hessian 的 LDL 分解对角线权重引入 Viterbi 分支度量,使搜索在编码块内跟随曲率。
正文
Abstract:The best two-bit weight quantizers for large language models, such as QTIP and Proteus, rotate each weight matrix by a random orthogonal transform, which must be undone at every decoding step, then encode it with a trellis or lattice code under a Euclidean search; the layer Hessian enters only through error feedback between coding blocks. We show that this leaves part of the Hessian unused. Error feedback turns the loss into a weighted sum of per-coordinate rounding errors whose weights, the diagonal of the Hessian's LDL factorization, existing quantizers compute but never read. We put these weights into the Viterbi branch metric, so the search follows the curvature within each coding block. This also explains the rotation: it removes this within-block variation, so weighting in the native basis and rotating are substitutes. On three models the weighted native search matches a full-dimension randomized Hadamard to within about one point of downstream accuracy, and weighting after the rotation gains little. Around this search we build CurveTQ, a trellis codec with no rotation, which handles the weights' amplitude and marginal shape with a factored scale field and a closed-form quantile table, and stores a start state per coding block so the trellis can adapt to the residual that error feedback carries into it. At two bits CurveTQ is 1-3 points higher in mean downstream accuracy than QTIP and Proteus on three 4-8B Instruct models, even after both are given our start state, which alone lifts either baseline by 1-3 points. It also leads on a 35B mixture of experts, to our knowledge the first trellis-coded result on such a model. With no rotation to undo, our decoder is the fastest of the three at all tested batch sizes and bit widths.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.09212 [cs.LG] |
| (or arXiv:2610.09212v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09212 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zi Wang [view email]
[v1]
Tue, 6 Oct 2026 23:14:06 UTC (175 KB)
来源:arXiv:cs.LG · arxiv.org