arXiv:cs.LG(机器学习,全量分类)· Alan Z. Song, Yinjie Chen, Mu Nan, Deva Ramanan, Michael J. Tarr, Andrew F. Luo·· 17 小时前AI 评分40
Vision Transformer 是否需要全对全注意力?VECA 用弹性学习核心实现全局通信
Do Vision Transformers Need All-to-All Attention? Global Communication Through Elastic Learned Cores
AI 导读
研究者提出 VECA(Visual Elastic-Core Attention),一种核心-边缘结构注意力的视觉 Transformer,patch token 仅通过少量学习核心交换全局信息,将注意力复杂度从 O(N²) 降至 O(N)。
正文
Abstract:Vision Transformers (ViTs) learn rich visual-semantic representations through all-to-all self-attention among patch tokens. However, this design implicitly assumes that direct pairwise patch interactions are necessary for effective representation learning. In this work, we challenge this assumption and show that representations supporting both global recognition and dense prediction can be learned without direct patch-to-patch interaction. We propose VECA (Visual Elastic-Core Attention), a vision transformer with core-periphery structured attention mediated by a small set of learned cores. Patch tokens exchange global information exclusively through these cores, while the full set of dense patches are preserved and iteratively updated across layers. This reduces attention complexity from $O(N^2)$ to $O(N)$, linear in the number of patches $N$ for a fixed core budget $C$. Unlike prior latent-token cross-attention architectures, VECA facilitates sparse global communication without compressing the spatial representation itself. Nested training along the core axis further enables a single model to elastically trade off computation and accuracy at inference time without retraining. Across image classification and dense prediction tasks, VECA remains competitive with full-attention backbones and outperforms the evaluated linear-complexity alternatives on most benchmarks. Moreover, without explicit supervision, these cores develop semantically organized structures that support object-label transfer across video frames. These results show that effective visual representations can be learned without direct all-to-all patch interaction.
| Comments: | Project repository here: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.12491 [cs.CV] |
| (or arXiv:2605.12491v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2605.12491 arXiv-issued DOI via DataCite |
Submission history
From: Zixi Song [view email]
[v1]
Tue, 12 May 2026 17:59:26 UTC (19,885 KB)
[v2]
Wed, 30 Sep 2026 19:03:36 UTC (28,248 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org