跳到正文
arXiv:cs.LG· Ruijie Li, Jiaxi Hu, Shiyu Wang, Yuxuan Liang·· 4 小时前AI 评分32

PHBA:前缀状态混合块注意力

PHBA: Prefix-State Hybrid Block Attention

AI 导读

研究提出 PHBA(Prefix-State Hybrid Block Attention),用 top-k 块稀疏检索替代 NHA 的固定局部滑窗注意力,并将每个检索块与其前缀状态配对。

正文

View PDF HTML (experimental)

Abstract:Hybrid architectures combining linear sequence models with softmax attention provide an effective balance between efficient long-context modeling and precise token retrieval. Existing designs such as Native Hybrid Attention (NHA) combine compressed long-term states with sliding-window attention, but their exact attention is restricted to a fixed local window. In this work, we introduce Prefix-State Hybrid Block Attention (PHBA), which replaces local sliding-window attention with top-k block-sparse retrieval and couples each retrieved block with a compact prefix state summarizing its preceding context. The prefix states are constructed by a gated linear recurrence at block boundaries and retrieved together with the corresponding token blocks, allowing the model to combine precise long-range evidence with compressed historical context within a unified layer. We further develop a hardware-aware Triton implementation that streams routed token blocks and prefix states without materializing large intermediate tensors. Experiments show that PHBA improves long-context and retrieval performance over strong linear and hybrid baselines while retaining efficient training and inference.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.08527 [cs.LG]
  (or arXiv:2610.08527v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.08527

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ruijie Li [view email]
[v1] Tue, 6 Oct 2026 15:22:49 UTC (1,015 KB)

来源:arXiv:cs.LG · arxiv.org