arXiv:cs.AI· Yuhan Chen, Siyuan Zhang, Nan Wang, Feiyang Kang, Ruoxi Jia·· 6 小时前AI 评分40
Hybrid Latent Attention:为循环语言模型压缩 KV cache 的混合潜在注意力机制
Hybrid Latent Attention for Looped Language Models
AI 导读
针对循环语言模型将 KV cache 扩大 T 倍的问题,研究者提出 Hybrid Latent Attention(HLA),在滑动窗口内保留精确的 key 和 value,窗口外的旧 token 则压缩为潜在表示供每个循环的 query 直接读取。
正文
Abstract:Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can decode at once and slows each decoding step, which reads the whole cache. We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values. We uptrain HLA on Ouro looped models (T=4) with 1.4B and 2.6B parameters, keeping the pretrained weights frozen and training only the added parameters to reproduce the original attention. The cache shrinks by 10.7x per token, fitting 4.0-8.8x as many concurrent sequences per GPU, and decoding throughput improves by 2.5x at 1K-token contexts and by up to 7.4x at 16K. HLA retains over 97% of the original accuracy on math, knowledge and reasoning benchmarks, and 96-100% on long-context retrieval up to 16K tokens. After supervised fine-tuning, it performs on par with the fine-tuned original model on competition-level math.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07940 [cs.CL] |
| (or arXiv:2610.07940v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07940 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yuhan Chen [view email]
[v1]
Tue, 6 Oct 2026 08:14:07 UTC (185 KB)
来源:arXiv:cs.AI · arxiv.org