arXiv:cs.LG(机器学习,全量分类)· Athanasios Zeris·· 14 小时前AI 评分30
FourierQK:滤波器形状、可容许性与泄漏-覆盖定律
FourierQK: Filter Shape, Admissibility and the Leakage-Coverage Law
AI 导读
FourierQK 通过受控消融实验发现,DC 与 Nyquist 分量对频率坍缩注意力有害(val ≈ 2.0),最优单尺度带宽为 σ ≈ 2 bins、中心在段落尺度(约 70 tokens),较 BASE-DOT 提升 Δ = +1.15 nats。
正文
Abstract:Frequency-collapse attention [Zeris, 2026e] achieves large gains over standard dot-product attention by replacing the Q/K dot product with a bandpass-filtered inner product at a learned frequency. A natural follow-up question is: which filter shape works best, and why? We test five hypotheses about filter properties -- DC suppression, Nyquist suppression, bandwidth, centre frequency, and multi-scale coverage -- using a controlled ablation on character-level language modelling (TinyShakespeare, 6-layer GPT). Our main findings are: (1) DC and Nyquist components are actively harmful (val ~= 2.0, equivalent to phase randomisation), confirming that oscillatory bandpass structure is essential, not just any low-dimensional spectral summary; (2) the optimal single-scale bandwidth is sigma ~= 2 bins centred at paragraph scale (~70 tokens), giving a clean gain of Delta = +1.15 nats over BASE-DOT; (3) admissible filters (zero-mean, Mexican Hat DOG m = 2) outperform non-admissible Gaussians at the same scale and provide partial protection against bilateral FFT leakage; (4) bilateral FFT leakage scales monotonically with spectral coverage -- narrowband filters (gap > +4) are clean, wideband filters (gap < +2) are leaky; and (5) causal time-domain Morlet at character scale cannot beat BASE-DOT (K=128 taps covers 50% of T=256 context), motivating word-level experiments in the companion MorletQK paper [Zeris, 2026f]. Together, findings (1)-(5) characterise FourierQK as effective in bidirectional attention settings (encoder-style, e.g. BERT), where full-sequence context is available at both training and inference time; autoregressive generation requires a causal spectral variant such as MorletQK [Zeris, 2026f] (decoder-style, e.g. GPT). Code available at: this https URL
| Comments: | 9 pages, 1 figure, 2 tables |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL); Signal Processing (eess.SP) |
| Cite as: | arXiv:2610.00009 [cs.LG] |
| (or arXiv:2610.00009v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00009 arXiv-issued DOI via DataCite |
Submission history
From: Athanasios Zeris [view email]
[v1]
Thu, 9 Jul 2026 14:41:04 UTC (315 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org