arXiv:cs.LG· Minsoo Cheong, Donghyun Son, Sungjoo Yoo·· 3 小时前AI 评分39
TaSQ:为 1-Bit KV Cache 压缩定制量化空间
Tailoring the Quantization Space for 1-Bit KV Cache Compression
AI 导读
TaSQ 通过查询引导的通道加权、跨头归一化和协方差感知的通道分组来定制向量量化目标空间,缓解现有 VQ 方法在 1-bit 极端压缩下的性能退化。该方法兼容 RoPE,可合并进投影权重与码本,保持传统 VQ 查表结构且服务开销可忽略。在单张 RTX 6000 Ada GPU 上,其 SGLang 实现支持最高 14 倍批大小,峰值吞吐量较 BF16 基线提升 1.87 倍。
正文
Abstract:The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To mitigate this bottleneck, vector quantization (VQ) has emerged as a promising approach for aggressive KV cache compression. However, existing VQ methods degrade substantially in the 1-bit regime. At such extreme compression, each codebook must represent a larger group of channels with a limited set of centroids, making effective use of its capacity increasingly challenging. To address this, we introduce $\textbf{TaSQ}$, which tailors the VQ target space by combining query-guided channel weighting, cross-head normalization, and covariance-aware channel grouping to better reflect the error sensitivity and statistical structure of cached activations. Since these transforms are RoPE-compatible and can be easily merged into projection weights and codebooks, TaSQ preserves the conventional VQ lookup structure and adds negligible serving overhead. Across general, long-chain-of-thought reasoning, and long-context retrieval benchmarks, TaSQ consistently outperforms existing low-bit KV cache VQ baselines while preserving reasoning stability. On a single RTX 6000 Ada GPU, its SGLang implementation supports up to $14\times$ larger batch sizes and achieves $1.87\times$ higher peak throughput compared to the BF16 baseline.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.03027 [cs.LG] |
| (or arXiv:2610.03027v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03027 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Minsoo Cheong [view email]
[v1]
Fri, 2 Oct 2026 09:04:25 UTC (672 KB)
来源:arXiv:cs.LG · arxiv.org