arXiv:cs.LG(机器学习,全量分类)· Haider Al-Tahan, Sean O'Brien, Anastasia Razdaibiedina, N. Apurva Ratan Murty·· 21 小时前AI 评分35
LAST:循环音频频谱图 Transformer
LAST: Looped Audio Spectrogram Transformer
AI 导读
Looped Audio Spectrogram Transformer(LAST)先处理全部 token,再复用同一组 block 仅对 class token 做多轮精炼,使后续处理开销很低。
正文
Abstract:Increasing depth of transformer models improves recognition, but it comes at a substantial cost. Each additional layer requires more parameters, which makes the process computationally inefficient. We ask whether additional processing can focus on integrating features already computed. Looped Audio Spectrogram Transformer (LAST) first processes all tokens, then reuses the same blocks to refine only the class token over fixed audio features, thereby making later passes inexpensive. On AudioSet, ten-pass LAST achieves 0.345 mean average precision, exceeding a twelve-layer sequential transformer by 2.1% relative with 49.4% fewer parameters, 42% fewer multiply-accumulate operations, and 9.8% higher measured throughput. Across separately trained models, increasing the pass count from two to ten improves accuracy while adding only 1.2% computation. Further evaluations show improved robustness to temporal masking and various other auditory augmentations, with better generalization on classification tasks with music, environmental, and event sounds.
| Comments: | 6 pages, 4 figures, 1 table |
| Subjects: | Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS) |
| ACM classes: | I.2.6; I.5.4; H.5.5 |
| Cite as: | arXiv:2610.01926 [cs.SD] |
| (or arXiv:2610.01926v1 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01926 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Haider Al-Tahan [view email]
[v1]
Thu, 1 Oct 2026 15:59:09 UTC (85 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org