arXiv:cs.LG· Xiaofan Que, Nir Elkayam, Spandan Pyakurel, Shuokai Pan, Dibakar Gope·· 2 天前AI 评分32
XOR-Trellis:超低复杂度反量化与曲率感知的无 Hadamard LLM 量化
XOR-Trellis: Ultra-Low-Complexity Dequantization and Curvature-Aware Hadamard-Free LLM Quantization
AI 导读
XOR-Trellis 提出两项互补技术,实现超低位宽 LLM 权重的格码量化:一是超低复杂度格码反量化器,采用结构化、硬件高效的状态到值映射,同时保留多样的重建选择以支持格码搜索;二是以曲率感知目标重新表述离散格码路径优化,在原始坐标空间中直接反映模型敏感度。两项技术结合,无需依赖基于 Hadamard 的不相干处理,即可实现高质量超低位格码量化,并支持低成本、高度并行的运行时重建。
正文
Abstract:Trellis-coded quantization enables high-dimensional compression of large language model (LLM) weights at ultra-low bit widths without the exponentially large codebooks required by conventional vector quantization. Practical deployment, however, presents two challenges: reconstructing compressed weights at sufficient parallel throughput to avoid making dequantization an inference bottleneck, and maintaining quantization accuracy without costly incoherence transformations. We address these challenges with two complementary techniques. First, we introduce an ultra-low-complexity trellis dequantizer that uses a structured, hardware-efficient state-to-value mapping while preserving diverse reconstruction choices for trellis search. Second, we reformulate discrete trellis path optimization with a curvature-aware objective that reflects model sensitivity directly in the original coordinate space. Together, these techniques enable high-quality ultra-low-bit trellis quantization with inexpensive, highly parallel runtime reconstruction and without relying on Hadamard-based incoherence processing.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.00432 [cs.LG] |
| (or arXiv:2610.00432v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00432 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Dibakar Gope [view email]
[v1]
Wed, 30 Sep 2026 17:05:40 UTC (4,604 KB)
来源:arXiv:cs.LG · arxiv.org