跳到正文
arXiv:cs.LG· Yixuan Liang·· 3 小时前AI 评分37

面向美国建筑电表数据的 AI 需求预测数据接入管线:设计、受控评估与一个被修正的负面结果

Validated Data Onboarding for AI Demand Forecasting on U.S. Building Meter Data: Design, Controlled Evaluation, and a Corrected Negative Result

AI 导读

该技术报告提出一套数据接入管线,在模型训练前检测并修复电表数据缺陷,且仅使用预测时可得的信息。在 Building Data Genome 2 的 12 栋美国建筑小时级电力数据(210。

正文

View PDF

Abstract:Electric utilities and grid operators increasingly rely on machine-learning models to forecast next-day demand, and those models learn from meter data that is routinely defective: readings go missing, sensors freeze, buildings read zero for hours, and units change by a factor of 100. This report presents a data-onboarding pipeline that detects and repairs such defects before a model is trained, using only information available at forecast time, and a controlled experiment that measures whether the pipeline protects a 24-hour-ahead forecast. On hourly electricity data for twelve U.S. buildings from the public Building Data Genome 2 dataset (210,528 rows, 2016-2017), seeded, hash-logged defects touching 0.10% of the training period raised the error of a gradient-boosting forecaster by 86%; after detection and past-only repair the error returned to the clean-data level (mean absolute scaled error 0.760 clean, 1.415 corrupted, 0.729 repaired) while 93% of training targets were retained. At a defect prevalence calibrated to published field studies (1.6% of training rows) the unprotected forecaster's error reached 4.4 times that of a seasonal-naive rule, and the repaired forecaster again matched the clean baseline. The same pattern held for ridge regression and a random forest and across horizons of 1 to 24 hours. A first version of the pipeline over-cleaned natural data and made forecasts 25% worse; that result is retained, its cause is traced in the published artifacts, and the per-building calibration that corrects it is documented as a dated amendment. Every number is reproducible from pinned public inputs with SHA-256 verification, 84 automated tests and continuous integration.
Comments: Technical report; 8 pages, 5 figures, 2 tables. Code, data manifests, and reproducibility artifacts: this https URL . Preprint; not peer reviewed
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02397 [cs.LG]
  (or arXiv:2610.02397v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02397

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yixuan Liang [view email]
[v1] Thu, 1 Oct 2026 19:23:37 UTC (819 KB)

来源:arXiv:cs.LG · arxiv.org