跳到正文
arXiv:cs.LG· Yuji Akamatsu, Takao Yamanaka·· 4 小时前AI 评分34

在基于注意力的时间序列预测中,注意力并不解释预测峰值:时间参考与预测输出的分化

When Attention Does Not Explain the Peak: Temporal Reference vs. Forecast Output in Attention-Based Time-Series Forecasting

AI 导读

一项针对Panama负荷数据集31个日对齐窗口的研究发现,基于注意力的负荷预测模型峰值时间误差中位数为0h、精确匹配率51.6%,而注意力描述符Ψ_out的argmax中位数误差为5h、精确匹配率为0%。

正文

View PDF HTML (experimental)

Abstract:Attention maps are often interpreted as evidence of what a forecasting model uses when making predictions. In our load-forecasting model, a CLS representation of historical demand queries 24 future exogenous horizon tokens through cross-attention, inviting a temporal interpretation in which highly attended horizons may appear to explain forecast peak timing. We test this interpretation using a horizon-level attention descriptor, $\Psi_{\mathrm{out}}$. Across 31 day-aligned windows of the Panama load dataset, the forecast achieves a median peak-time error of 0 h and a 51.6% exact-match rate, whereas the argmax of $\Psi_{\mathrm{out}}$ has a median error of 5 h and 0% exact match. The forecast peak is closer to the observed peak in 27 of 31 windows. This dissociation is not merely an argmax artifact: within $\pm1$ h, attention reaches only $1.16\times$, $1.11\times$, and $1.14\times$ the uniform baseline around observed, predicted, and weekly-naive peaks, respectively, indicating weak and non-selective concentration. Yet the attention profile is structured, with cross-window consistency of 0.83. Replacing 12 future weather features with their training-set means makes the profile nearly uniform, showing sensitivity to future weather variation rather than fixed horizon position alone. The dissociation is also reproduced across three random-seed runs. These results show that structured, input-sensitive, and reproducible horizon-level cross-attention need not provide a valid peak-selective explanation of forecast behavior. The observed behavior is instead consistent with an internal horizon-reference role for integrating future exogenous information, although this functional role is not causally established.
Comments: NeurIPS 2026: Accepted to the TAE (Trust-AI-Eval) Workshop
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.07080 [cs.LG]
  (or arXiv:2610.07080v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07080

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yuji Akamatsu [view email]
[v1] Mon, 5 Oct 2026 11:15:35 UTC (475 KB)

来源:arXiv:cs.LG · arxiv.org