跳到正文
arXiv:cs.LG· Heejun Kim, Junyoung Lee, SangLyul Cho, Dongsu Han, Insu Han, Sehoon Kim·· 4 小时前AI 评分33

ResidualQuant:面向循环 Transformer 的 2-Bit 残差 KV Cache 量化

ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals

AI 导读

针对循环 Transformer 中 KV cache 随循环次数增长的内存瓶颈,ResidualQuant 以最后一轮循环的 KV 状态为参考,用低精度残差表示其余循环,并结合最小二乘缩放、旋转与逐循环混合精度,实现低至 INT2 的量化。

正文

View PDF HTML (experimental)

Abstract:Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count. However, KV cache memory still scales with the number of loops, becoming a key memory bottleneck that limits batch size and inference throughput. KV cache quantization can alleviate this bottleneck, but existing methods often suffer substantial accuracy degradation at aggressive low-precision regimes. We observe that looped Transformers offer a unique opportunity: KV states across loops are highly similar. Based on this observation, we propose ResidualQuant, which uses the final-loop KV states as a reference and represents the remaining loops with low-precision residuals. Our method further combines least-square scaling and rotations applied to the residuals, as well as loop-wise mixed precision, to enable accurate quantization down to INT2 while retaining efficient reconstruction. Across multiple looped Transformer models and mathematical reasoning and code generation benchmarks, ResidualQuant consistently improves the accuracy-memory tradeoff over state-of-the-art rotation-based KV quantization. In particular, our method retains accuracy close to BF16 under mixed-precision settings while reducing theoretical KV storage by 80.7%, achieving up to 13.0% higher accuracy than the rotation-based baseline at the same memory budget. On an RTX 5090, the reduced KV memory traffic improves fixed-batch decode throughput by up to 2.73x, while the smaller memory footprint enables up to 2x larger batches, improving peak throughput by up to 4.15x.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.10381 [cs.LG]
  (or arXiv:2610.10381v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.10381

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sehoon Kim [view email]
[v1] Wed, 7 Oct 2026 16:40:21 UTC (1,022 KB)

来源:arXiv:cs.LG · arxiv.org