跳到正文
arXiv:cs.LG· Roberto Stanzione, Jules Barbe, Magali Parrino, J\'er\'emie Fourmann, Paul Boniol·· 2 天前AI 评分35

SHAD:面向时间序列异常检测、可解释性与可解读性的端到端基准

Detect, Explain, Interpret: An End-to-End Benchmark for Time Series Anomaly Detection, Explainability and Interpretability

AI 导读

研究者发布 SHAD 基准,包含 215 条来自 Scality 真实分布式云存储系统的高维多变量时间序列,覆盖三类不同严重程度的异常。该基准同时评测检测、可解释性与可解读性,并给出覆盖 TSAD 全流程的基线结果,包括用冻结 LLM 基线定位和解读异常。

正文

View PDF HTML (experimental)

Abstract:Time Series Anomaly Detection has received increasing attention, driven by the growing availability of complex time series data. This surge has led to the development of numerous detection methods, as well as a variety of benchmarks aimed at thoroughly evaluating their performance. However, most existing detectors remain largely agnostic to domain context, overlooking explainability and interpretability. One of the main reasons for this gap is that current benchmarks primarily focus on detection accuracy, and only few of them evaluate spatial explainability. Moreover, no benchmark currently provides sufficiently rich semantic annotations to support the generation of human-understandable interpretations of anomalies. To address these limitations, we introduce SHAD (Scality High-dimensional Anomaly Detection benchmark), a fully annotated benchmark composed of 215 multivariate, high-dimensional time series collected from real-world distributed cloud storage systems operated by Scality. The proposed dataset includes rich contextual information, covering three families of anomalies with varying degrees of severity. As further contribution, we provide a foundation for future work by evaluating baseline methods for Detection, Explainability, and Interpretability, covering all stages of a TSAD pipeline. For Detection, we benchmark a wide range of existing anomaly detectors, testing their effectiveness on the proposed real-world dataset. Then, we consider explainability by evaluating whether measuring the contribution of each dimension in the generated anomaly score can provide accurate anomaly attributions. Finally, for interpretability, we investigate the effectiveness of frozen LLM baselines in localizing and interpreting anomalies.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Databases (cs.DB)
Cite as: arXiv:2610.01168 [cs.LG]
  (or arXiv:2610.01168v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01168

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Paul Boniol [view email]
[v1] Thu, 1 Oct 2026 06:44:55 UTC (7,265 KB)

来源:arXiv:cs.LG · arxiv.org