跳到正文
arXiv:cs.LG· Amit Rajula·· 4 小时前AI 评分39

基于滞后数据快照训练的模型差分刷新策略:从单一年龄等价极限到最优分段分配

Differential Refresh Policies for Models Trained on Lagging Data Snapshots: From a Single-Age Equivalence Limit to an Optimal Per-Segment Allocation

AI 导读

研究发现,任何基于单一全局训练数据年龄的静态单调刷新触发器,在运行上等价于校准过的统一年龄计时器,全局陈旧度评分不携带时钟之外的调度信息。改用按段差分刷新、各段独立设定年龄与刷新间隔后,每段最优刷新率正比于其风险 w_j λ_j 的平方根,成本不超过统一计时器。

正文

View PDF HTML (experimental)

Abstract:Production machine-learning models are derived artifacts of time-bounded training snapshots: a deployed model is a materialized view over a training cut that ages the instant it is built. A common response is to replace the fixed retraining cadence with an adaptive trigger -- a weighted staleness score that retrains when accumulated source risk crosses a threshold. We show this is the wrong lever, and identify the right one. First, an equivalence limit: any refresh trigger that is a static, strictly monotone function of a single shared global training-data age is operationally equivalent to a calibrated uniform age timer, so a global staleness budget, however elaborately it weights segments, sources, and sensitivities, carries no scheduling information a clock does not. The limit also shows how to escape it: refresh segments differentially, giving each its own age and refresh interval, which is meaningful when refresh cost is separable across segments (incremental training or per-segment models). We solve the resulting budget-allocation problem. In the frequent-refresh regime each segment's optimal refresh rate is proportional to the square root of its risk $w_j \lambda_j$ (weight times change rate), and the optimal policy never costs more than the uniform timer, beating it by a closed-form Cauchy-Schwarz "price of uniformity" that is zero for homogeneous workloads and grows with heterogeneity. In a discrete-event simulation with real Poisson change events, the optimal policy lowers realized weighted stale exposure by 8-29% relative to the uniform timer at matched refresh budget, winning on 86-100% of seeds; a naive exposure-threshold policy does not, showing the allocation is what helps; and the advantage survives 50% rate-estimation noise. The leverage in model refresh is not a better score but a better action.
Subjects: Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC)
Cite as: arXiv:2610.09519 [cs.LG]
  (or arXiv:2610.09519v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09519

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Amit Rajula [view email]
[v1] Wed, 7 Oct 2026 06:15:40 UTC (41 KB)

来源:arXiv:cs.LG · arxiv.org