跳到正文
arXiv:cs.CL· Xu Yang, Jiapeng Zhang, Yuxin Chen, Feiqiang Sun, Chengguang Xu, Feng Jin, Zhuo Tang·· 4 小时前AI 评分39

Self-Indexing Attention:面向压缩兼容稀疏长上下文 LLM 推理的无训练框架

Self-Indexing Attention for Compression-Compatible Sparse Long-Context LLM Inference

AI 导读

Self-Indexing Attention 是一个基于共享变换域符号-幅值表示的无训练框架,用可复用 token 级 1-bit 索引统一 prefill 分组选择与 decode 检索,并与外部 KV-cache 压缩兼容而无需额外索引元数据。

正文

View PDF HTML (experimental)

Abstract:Sparse long-context inference requires efficient token retrieval in both prefill and decode. Existing methods often use different retrieval strategies for the two stages, preventing one retrieval representation from being reused throughout inference. We propose Self-Indexing Attention, a training-free framework built on a shared transform-domain sign-magnitude representation. The key signs provide a reusable token-level index for grouped prefill selection and decode retrieval, while the same representation remains compatible with external KV-cache compression without separate indexer metadata. This 1-bit index enables efficient retrieval through bitwise operations widely supported by modern accelerators. At 5% attention density, Self-Indexing Attention remains close to dense attention on LongBench and RULER and achieves up to 6.1x prefill and 10.3x decode attention-operator speedups. Experiments with TurboQuant and DeepSeekV4-Flash further demonstrate compatibility with low-bit KV-cache compression and pretrained sparse-attention indexers.
Subjects: Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2609.13205 [cs.IR]
  (or arXiv:2609.13205v2 [cs.IR] for this version)
  https://doi.org/10.48550/arXiv.2609.13205

arXiv-issued DOI via DataCite

Submission history

From: Xu Yang [view email]
[v1] Sun, 16 Aug 2026 03:41:37 UTC (214 KB)
[v2] Wed, 7 Oct 2026 07:41:29 UTC (472 KB)

来源:arXiv:cs.CL · arxiv.org