跳到正文
arXiv:cs.AI· Yuhan Chen, Siyuan Zhang, Nan Wang, Feiyang Kang, Ruoxi Jia·· 6 小时前AI 评分40

Hybrid Latent Attention:为循环语言模型压缩 KV cache 的混合潜在注意力机制

Hybrid Latent Attention for Looped Language Models

AI 导读

针对循环语言模型将 KV cache 扩大 T 倍的问题,研究者提出 Hybrid Latent Attention(HLA),在滑动窗口内保留精确的 key 和 value,窗口外的旧 token 则压缩为潜在表示供每个循环的 query 直接读取。

正文

View PDF HTML (experimental)

Abstract:Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can decode at once and slows each decoding step, which reads the whole cache. We propose Hybrid Latent Attention (HLA), which keeps exact keys and values within a sliding window of W recent tokens and stores each older token as a compact latent that the query of each loop reads directly, without reconstructing keys and values. We uptrain HLA on Ouro looped models (T=4) with 1.4B and 2.6B parameters, keeping the pretrained weights frozen and training only the added parameters to reproduce the original attention. The cache shrinks by 10.7x per token, fitting 4.0-8.8x as many concurrent sequences per GPU, and decoding throughput improves by 2.5x at 1K-token contexts and by up to 7.4x at 16K. HLA retains over 97% of the original accuracy on math, knowledge and reasoning benchmarks, and 96-100% on long-context retrieval up to 16K tokens. After supervised fine-tuning, it performs on par with the fine-tuned original model on competition-level math.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07940 [cs.CL]
  (or arXiv:2610.07940v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.07940

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yuhan Chen [view email]
[v1] Tue, 6 Oct 2026 08:14:07 UTC (185 KB)

来源:arXiv:cs.AI · arxiv.org