跳到正文
arXiv:cs.LG· Joe McKenna, Anastasios Alexandridis, Nathan Susanj, Jing Liu·· 2 天前AI 评分36

CommunityKV:通过图划分实现高效长上下文解码

CommunityKV: Efficient Long-Context Decoding via Graph Partitioning

AI 导读

CommunityKV 将稀疏注意力建模为社区检测问题,利用标准 prefill 阶段已算出的 QK^T 分数构建 token 图并划分为社区,以检索语义连贯的 token 组;其局部更新规则可在常数时间内将新生成 token 分配到社区,无需全局重新划分。

正文

View PDF HTML (experimental)

Abstract:Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the linear growth of key-value cache memory transfer. Sparse attention mitigates this by retrieving only relevant tokens, but current approaches either require large-scale training or, within the training-free regime, rely on semantically coarse heuristics or expensive clustering that is difficult to update efficiently during decoding. We introduce CommunityKV, a framework that formulates sparse attention as a community detection problem. CommunityKV constructs a token graph from the $QK^T$ scores already computed during standard prefill, and partitions the graph into communities to enable retrieval of semantically coherent token groups. A local update rule assigns newly generated tokens to communities in constant time, enabling sparse retrieval throughout streaming decoding without global re-partitioning. We evaluate CommunityKV on Qwen3 and Llama-3.1 models across three long-context benchmarks. With one graph per query head, CommunityKV delivers up to $1.25\times$ the end-to-end generation throughput of dense attention, while query-group graph aggregation yields up to $1.71\times$ with comparable accuracy.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.00418 [cs.LG]
  (or arXiv:2610.00418v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00418

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Joe McKenna [view email]
[v1] Wed, 30 Sep 2026 14:36:50 UTC (935 KB)

来源:arXiv:cs.LG · arxiv.org