arXiv:cs.LG· Kailen Hargenrader, Edoardo Calvello, Bohan Chen·· 2 天前AI 评分34
重尾测度映射学习中的注意力核:post-norm Transformer 的慢增长核设计
Attention Kernels for Learning Maps Between Heavy-Tailed Measures
AI 导读
针对多项式尾概率测度上算子学习时 softmax 指数加权导致测度级注意力积分发散的问题,研究用增长更慢的函数替换指数,并构建了两个带闭式目标的基准。在无数据变换时,softmax 模型在两个重尾基准上出现集成坍缩,而三种慢增长核均避免坍缩;Symlog 预处理仅让 softmax 在矩阵求逆任务上避免坍缩,在 sheared swap 任务上仍失败。
正文
Abstract:Operator learning on probability measures can be accomplished with transformers. For measures with polynomial tails, the exponential weighting in softmax can make the corresponding measure-level attention integrals diverge. This motivates replacing the exponential with slower-growing functions. We construct two benchmarks for operator learning on measures with closed-form targets. We use these benchmarks to study attention kernel growth and data transformation in post-norm transformers. Without data transformation, the softmax models exhibit ensemble collapse on both heavy-tailed benchmarks, while the three slower-growing kernels avoid collapse. Symlog preprocessing allows softmax to avoid collapse on the matrix inverse task but not on the sheared swap task. On the Gaussian control, all four kernels perform similarly. We also examine how sample size affects the sensitivity of empirical energy and Wasserstein distances to tail differences. These results support slower-growing attention kernels as an effective design choice for post-norm transformers learning from heavy-tailed ensembles.
| Comments: | 32 pages, 17 figures, accepted to NeurIPS 2026 Workshop on AI for Stochastic Dynamics |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00564 [cs.LG] |
| (or arXiv:2610.00564v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00564 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kailen Hargenrader [view email]
[v1]
Wed, 30 Sep 2026 18:38:24 UTC (1,024 KB)
来源:arXiv:cs.LG · arxiv.org