arXiv:cs.AI· Yuyang Zhao, Lian Xu, Hao Xue·· 5 小时前AI 评分32
时间序列预测基准需要基于场景的压力测试
Time Series Forecasting Benchmarks Need Scenario-Grounded Stress Testing
AI 导读
研究者提出时间序列预测(TSF)评测应转向"场景化压力测试",每个测试实例需包含历史输入、未来目标、语义场景、显式失效算子和可量化难度等级。现有基准只奖励低留出误差,鲁棒性研究也仅用高斯噪声、随机掩码或有界对抗扰动模拟失效,无法反映部署系统的真实失效模式。TSF 基础模型的兴起使这一评测缺口更紧迫,因为不可审计的预训练语料让留出集泛化越来越不可靠。
正文
Abstract:Time series forecasting (TSF) increasingly drives decisions in transportation, energy, finance, healthcare, and infrastructure, yet current evaluation remains overly narrow: standard benchmarks reward low held-out error, while robustness studies typically reduce failure to Gaussian noise, random masking, or bounded adversarial perturbations. This obscures the real failure modes of deployed forecasting systems. Input-side anomalies are not merely noisier inputs: they often reflect structured events that alter temporal dynamics, break cross-variable dependencies, induce regime shifts, or propagate from faulty sensors to downstream decisions. These semantic, causal, and system-level failures cannot be faithfully captured by i.i.d. perturbations alone. The rise of TSF foundation models makes this evaluation gap more urgent, as unauditable pretraining corpora make held-out generalization increasingly unreliable. We therefore advocate scenario-grounded stress testing. Each test instance should include historical inputs and future targets, together with a semantic scenario, an explicit failure operator, and a measurable difficulty level. This shift makes evaluation interpretable, attributable, and deployment-relevant and friendly, enabling the community to ask not only which model is accurate, but under what conditions it fails and why.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02608 [cs.AI] |
| (or arXiv:2610.02608v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02608 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hao Xue [view email]
[v1]
Fri, 2 Oct 2026 00:01:12 UTC (778 KB)
来源:arXiv:cs.AI · arxiv.org