跳到正文
arXiv:cs.LG· Weijia Han, Lisha Qu, Zhenda Li, Liying Liang·· 4 小时前AI 评分42

自监控测试时自适应中的可交换性失效问题

The Premise Is the Problem: Exchangeability Failure in Self-Monitored Test-Time Adaptation

AI 导读

一项针对多步时间序列预测的研究发现,当监控与自适应使用同一预测误差反馈时,重叠目标和预测误差中的依赖性会破坏监控器统计保证所需的关键假设。监控器会在没有有害变化时误报,其响应反而损害预测质量;自适应还会向自身监控器隐藏持续变化,而原始冻结模型保留了更清晰的信号。

正文

View PDF HTML (experimental)

Abstract:Modern forecasting models are often updated after deployment so they can respond to changing data. These updates can also make predictions worse, so practical systems need a reliable monitor that can detect harmful changes and trigger protection. A natural design is to monitor the same prediction errors that guide the updates. This paper asks whether the statistical guarantee behind such a monitor remains valid when monitoring and adaptation use the same feedback. We study this question in multi-step time-series forecasting. We show that overlapping targets and dependence in forecast errors can break a key assumption required by the guarantee. The monitor may then raise alarms even when no harmful change has occurred, and its response can further damage prediction quality. We also find that adaptation can hide sustained changes from its own monitor, while the original frozen model retains a clearer signal. These results expose a basic failure mode in self-monitored adaptation. They show why reliable deployment requires checking the monitor's assumptions, comparing adaptation with the frozen model under realistic feedback, and limiting the effect of every protective response.
Comments: 49 pages, 9 figures
Subjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as: arXiv:2610.07038 [cs.LG]
  (or arXiv:2610.07038v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07038

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Weijia Han [view email]
[v1] Sun, 4 Oct 2026 23:17:08 UTC (514 KB)

来源:arXiv:cs.LG · arxiv.org