arXiv:cs.LG· Zong-Han Bai, Po-Yen Chu·· 5 小时前AI 评分35
间歇性需求预测中池化何时有效?遗忘机制下的可信度与分辨率研究
When Does Pooling Pay? Credibility and Resolution under Forgetting in Intermittent-Demand Forecasting
AI 导读
研究提出拟合窗口诊断方法,在分层经验贝叶斯 hurdle 模型中统一了序列自身历史保留与跨序列借力的决策,形式化了可信度边界与分辨率条件。在五个公开间歇性需求面板(超19000条序列)上,诊断在两个方向上均判断正确:截断降低可信度时共享先验有效,关闭它在最短历史上代价为5%和12%,而全长度仅约1%。所得预测器 EBB 在固定起点下与逐序列高斯过程专家差距约3%以内,成本低两到三个数量级。
正文
Abstract:Forecasting many sparse series requires two choices: how much of each series' own past to retain, and how much to borrow from other series. We show that when one exponential recency operator is applied to item- and group-level statistics alike in a hierarchical empirical-Bayes hurdle model, the two become one decision: forgetting preserves the leverage of a shared prior while raising the noise floor against which fine cross-series structure must be resolved. We formalize this two-sided effect as a fitting-window diagnostic, a credibility bound and a resolution condition for candidate pools, and derive from it an ex ante screen for regimes where refinement cannot pay. On five public intermittent-demand panels comprising over 19,000 series, where item-level evidence is sparsest, the diagnostic is right in both directions. Where truncation lowers credibility the shared prior pays: switching it off costs $5\%$ and $12\%$ at the shortest histories on the two panels it helps most, against about $1\%$ at full length, and the gain sits in the occurrence block. On four panels the diagnostic reports at most $0.6\%$ of room for refinement and realized gains are no larger; on the one panel with room the learned mixture does not collect it, which locates the open problem in the partition objective. The resulting forecaster, EBB, stays within about $3\%$ of a per-series Gaussian-process specialist at fixed origin and within $1\%$ under walk-forward evaluation where both run, at two to three orders of magnitude lower cost.
| Comments: | Preprint. 52 pages, 15 figures. Equal contribution by the two authors. v3: Retitled and substantially revised; reframes the work around forgetting and pooling, with credibility/resolution theory and expanded experiments |
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG) |
| Cite as: | arXiv:2511.12749 [stat.ML] |
| (or arXiv:2511.12749v3 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2511.12749 arXiv-issued DOI via DataCite |
Submission history
From: Po-Yen Chu [view email]
[v1]
Sun, 16 Nov 2025 19:24:33 UTC (137 KB)
[v2]
Wed, 1 Apr 2026 14:48:04 UTC (594 KB)
[v3]
Fri, 2 Oct 2026 03:18:06 UTC (884 KB)
来源:arXiv:cs.LG · arxiv.org