arXiv:cs.LG· Mohamed Readh Fentazi, Mazene Ameur, Adlen Ksentini·· 4 小时前
时间序列预测的测试时适配延迟审计:标签延迟如何影响 TTA 效果
Training on the Future: A Delay-Aware Audit of Test-Time Adaptation for Time-Series Forecasting
AI 导读
研究团队构建了一套无泄漏测试框架,强制预测起点 s 的标签仅在 s+d(d≥H)步后释放,并对 TAFAS、COSA、PETSA、DynaTTA 四种近期测试时适配方法在 ETTm1、ETTh2、Weather、Electricity、Traffic 五个基准上做因果延迟审计。
正文
Abstract:Test-time adaptation (TTA) methods for time-series forecasting update a deployed model, or a small adapter around it, from incoming ground truth. But the label of an $H$-step forecast exists only $H$ steps later, and real data pipelines add further delay. We build a leakage-free harness in which the label of forecast origin $s$ is released for updates only at step $s+d$ with $d \ge H$, and enforce this rule inside the released code of four recent TTA methods (TAFAS, COSA, PETSA and DynaTTA), run on their own backbones and checkpoints across five benchmarks (ETTm1, ETTh2, Weather, Electricity and Traffic). As references we add two closed-form correctors: a bank of recursive least squares (RLS) filters combined by a per-coordinate median, with no tunable hyperparameters and 56 microseconds per step on the 7-channel streams, and an ELF-style linear corrector. Under causal delayed labels the picture is asymmetric. On ETTm1 every audited method genuinely adapts, yet the RLS bank still beats three of the four at a fraction of their cost; only DynaTTA beats the bank, only at the minimum causal delay, and at roughly 2,500 times the per-update cost; the ELF-style corrector beats all four. On the other four datasets, the largest statistically significant improvement any published method achieves over its own frozen checkpoint is half a percent, on all four at least one published method is significantly worse than the frozen model at the minimum causal delay, and on drift-heavy ETTh2 longer label delays make every adapter that separates from the frozen model, ours included, significantly harmful. Leaky next-step updates inflate the apparent gains of simple adapters by up to 110%, and the backbone training recipe moves frozen online error by up to a factor of 25, more than any adaptation effect we measure. We release the harness, integration patches and all cached runs.
| Comments: | 34 pages, 9 figures. Code and cached results: this https URL |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.12232 [cs.LG] |
| (or arXiv:2610.12232v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12232 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mohamed Readh Fentazi [view email]
[v1]
Thu, 8 Oct 2026 16:16:58 UTC (15,775 KB)
来源:arXiv:cs.LG · arxiv.org