跳到正文
arXiv:cs.LG· Zhida Wang, Wei Lyu, Guo Yu, Sui Tang·· 4 小时前

SPD-MetaFormer:小数据脑电解码的无注意力 SPD 架构

SPD-MetaFormer is what you need for small-data brain decoding

AI 导读

研究提出 SPD-MetaFormer,一种基于 log-Euclidean 几何下均匀加权 Fréchet 聚合的无注意力架构,用于小数据脑信号解码。作者发现 MAtt 与 GBWAtt 训练后的注意力权重接近均匀,替换为均匀权重对平均预测性能影响甚微;SPD-MetaFormer 在三个 EEG 基准上取得了与已发表的欧氏及流形基线相当的结果。

正文

View PDF HTML (experimental)

Abstract:Brain signal decoding is challenging because neural recordings are noisy and vary across individuals, while labeled data are often limited. Recent attention-based models on the symmetric positive definite (SPD) manifold have nevertheless achieved strong performance using covariance and connectivity representations, yet the contribution of learned token weighting remains unclear. We examine two representative architectures, MAtt (based on log-Euclidean geometry) and GBWAtt (based on generalized Bures--Wasserstein geometry), and find that their learned attention weights remain close to uniform after training. We relate this behavior to bounded similarity parameterizations that, under the original softmax scaling, limit attention-weight contrast. Moreover, replacing learned weights with uniform weights, throughout training and evaluation, has little effect on mean predictive performance while preserving each model's original aggregation geometry. Motivated by these findings, we introduce SPD-MetaFormer, an attention-free architecture built on uniformly weighted Fréchet aggregation under log-Euclidean geometry. Its backbone uses a geodesic residual to update a summary token and a shared spectral feedforward map to transform all tokens, followed by a learned weighted readout. Token states remain SPD-valued until tangent-space classification. Across three EEG benchmarks, SPD-MetaFormer achieves competitive results relative to published Euclidean and manifold baselines. Separate matched reproductions test learned versus uniform weighting within MAtt and GBWAtt. These results suggest that, in the short-sequence and limited-data regimes studied, carefully designed SPD architectures can provide a simpler and effective alternative to adaptive manifold attention.
Subjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as: arXiv:2610.10952 [cs.LG]
  (or arXiv:2610.10952v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.10952

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sui Tang [view email]
[v1] Wed, 7 Oct 2026 22:09:32 UTC (378 KB)

来源:arXiv:cs.LG · arxiv.org