arXiv:cs.LG· Martin Hofmann, Patrick M\"ader·· 4 小时前AI 评分33
网络训练历史何时比当前状态更能预测未来学习?来自响应探针与预测筛选的证据
When does a network's training history predict its future learning better than its current state? Evidence from a response probe and a forecasting screen
AI 导读
研究对比了小型多层感知机的训练历史与当前状态对未来学习的预测能力:在42种历史轨迹、4个检查点上,至多四维的历史状态未能优于校准后的当前状态模型(增益-21.4%,90%区间[-91.9, 8.1])。
正文
Abstract:Networks that behave alike now can still learn differently when training continues. Work on loss of plasticity and critical periods shows that the path to a state shapes what follows; it does not show whether the path carries information that a measurement of the state itself misses. We ask when the training history of a network predicts its future learning better than its current state. In a main study, small multilayer perceptrons were trained under three history regimes (42 histories), and future learning was measured at four checkpoints by a short probe: a copy of the network trained for 100 updates on a new task. Before the prediction result was read, the protocol checked the probe. It responded monotonically to a function-preserving rescaling of hidden units, repeated measurements agreed (intraclass correlation 0.940, [0.903, 0.997], in the least reliable class, mean of three repeats), and a re-initialisation of units was visible directly after it but not 100 to 200 updates later. A history state of at most four dimensions did not improve on a calibrated model of the current state (gain -21.4%, 90% interval [-91.9, 8.1]; required in advance: 10%). A companion screen on 1,560 synthetic regression runs asked the same question for a target further away, the final error of the run. There, history models forecast better than the current validation error after 12 of up to 240 epochs (compact state 30.3%, [15.8, 39.4], a contextual comparison) and were not distinguishable from it after 48. In both studies the history was informative only while the current state was not yet informative about the target; this reading was formed after the results.
| Comments: | 18 pages, 6 figures, 5 tables |
| Subjects: | Machine Learning (cs.LG) |
| MSC classes: | 68T05 |
| ACM classes: | I.2.6 |
| Cite as: | arXiv:2610.09621 [cs.LG] |
| (or arXiv:2610.09621v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09621 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Martin Hofmann [view email]
[v1]
Wed, 7 Oct 2026 07:59:53 UTC (141 KB)
来源:arXiv:cs.LG · arxiv.org