arXiv:cs.LG(机器学习,全量分类)· Michael Fore, James Mason Inder, Mrishika Nair, Praneetha Vaddamanu, Sharlina Keshava·· 14 小时前AI 评分34
研究拆解 Chronos-2 组注意力:池化有益,学习权重在上下文学习中反而有害
Pooling Helps, Learned Weighting Hurts In-Context: Decomposing Group Attention
AI 导读
研究拆解 Chronos-2 提出的组注意力机制,发现统一池化(仅 V/O、无 Q/K 加权)在 20 个传感器网络配置中有 18 个带来正向收益,而学习到的 Q/K 加权在多变量预测中正向或可忽略,却显著拖累 10 个传感器网络 ICL 配置中的 8 个,其中 4 个甚至差于单变量推理。仅将第一个 block 的注意力矩阵 α 统一化,即可改善所有测试的 ICL 配置。
正文
Abstract:Group attention, introduced by the time series forecasting model Chronos-2, attends over the variates of a group at a fixed patch index and serves both multivariate (MV) and in-context learning (ICL) forecasting. Rather than evaluating this cross-variate attention design as a whole, we ask which part of the mechanism earns the benefit and probe its applicability to both MV and ICL regimes. By editing the attention matrix $\alpha$ at inference we separate the two pathways a head comprises: V/O, which projects a weighted summary of the group, and Q/K, which decides the weights. Uniform pooling (V/O without any Q/K weighting) is positive on 18 of our 20 sensor-network configurations, while the learned weighting (Q/K) splits by group type: its contribution is positive or negligible for MV, but materially degrades 8 of the 10 sensor-network ICL configurations, leaving 4 of them worse than univariate inference. By isolating the impact of different layers, we find that uniforming $\alpha$ in the first block alone improves every ICL configuration we test.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.01831 [cs.LG] |
| (or arXiv:2610.01831v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01831 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Michael Fore [view email]
[v1]
Thu, 1 Oct 2026 15:05:41 UTC (94 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org