跳到正文
arXiv:cs.LG· Ryan Solgi, Parsa Madinei, Zheng Zhang·· 4 小时前AI 评分35

TuBA:面向多维序列建模的 Tucker 瓶颈注意力机制

Tucker Bottleneck Attention for Multi-Dimensional Sequence Modeling

AI 导读

研究者提出 Tucker 瓶颈注意力(TuBA),利用低秩张量结构将隐藏张量投影到紧凑 Tucker 核心上做多头自注意力,实现次二次复杂度计算。在视频预测和全球天气预报任务上,TuBA 相比标准自注意力最多降低 24.7% 和 37.1% 的误差,计算量分别减少 66.6% 和 85.1%,最高提速 4.27 倍。其自回归版本在核心内做双向交互、核心间做因果注意力。

正文

View PDF HTML (experimental)

Abstract:The quadratic cost of self-attention limits scalability to long sequences from multidimensional data. We introduce Tucker bottleneck attention (TuBA), which exploits low-rank tensor structure for efficient global token mixing. TuBA projects hidden tensors into compact Tucker cores, performs multi-head self-attention and linear projections on the cores, and writes updates back to the ambient space, enabling subquadratic computation. Its autoregressive extension combines bidirectional interactions within cores with causal attention across cores. On video prediction and global weather forecasting, TuBA achieves favorable accuracy-efficiency trade-offs over standard and efficient attention and task-specific models. Compared to standard self-attention, TuBA reduces error and computation by up to 24.7% and 66.6% for video prediction and 37.1% and 85.1% for autoregressive weather forecasting, with speedups up to 4.27 times. Low-rank Tucker cores and multi-frame generation also outperform full-rank attention and frame-by-frame generation, respectively.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.09090 [cs.LG]
  (or arXiv:2610.09090v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09090

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ryan Solgi [view email]
[v1] Tue, 6 Oct 2026 20:45:21 UTC (835 KB)

来源:arXiv:cs.LG · arxiv.org