跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Riccardo Ali, Alessio Borgi, Mario Severino, Alessio Gravina, Davide Bacciu, Pietro Li\`o, Christopher Irwin·· 14 小时前AI 评分34

让注意力头对话:超越对角图的注意力机制

Let the Heads Talk: Beyond Diagonal Graph Attention

AI 导读

研究者提出 Topological Attention(Top-A),一种允许注意力头之间进行边条件跨头通信的多头注意力机制,突破了标准多头注意力只能实现对角边映射的限制。

正文

View PDF HTML (experimental)

Abstract:Sheaf Neural Networks generalize scalar-weighted message passing by replacing scalar edge weights with linear transport maps between local feature spaces. Yet the role of this matrix-valued transport is entangled with the broader sheaf-diffusion construction. We isolate the transport primitive through quiver representations and establish a direct connection with multi-head attention. Treating attention heads as coordinates of a local transport space reveals that standard multi-head attention implements diagonal edge maps: along each directed interaction, a source head can contribute only to the corresponding receiver head. Allowing off-diagonal entries instead enables edge-conditioned communication across heads before neighborhood aggregation. We show that this operation cannot, in general, be absorbed into a single shared linear map applied after aggregation. Building on this characterization, we introduce Topological Attention (Top-A), a multi-head attention that learns edge-dependent off-diagonal routes while preserving the original same-head paths and exactly recovering vanilla attention when the additional routing vanishes. We evaluate Top-A on relational reasoning, heterogeneous graph learning, and algorithmic reasoning, including out-of-distribution generalization, with heterophilic node classification as a contrast setting. The results show that cross-head transport is most useful when the task benefits from interaction-dependent transformations, while heterophily alone provides no systematic advantage. These findings identify edge-conditioned cross-head communication as a distinct computational primitive of matrix-valued transport.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.01494 [cs.LG]
  (or arXiv:2610.01494v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01494

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Alessio Borgi Dr. [view email]
[v1] Thu, 1 Oct 2026 11:37:18 UTC (2,009 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org