arXiv:cs.LG(机器学习,全量分类)· Yuxin Ma, Adir Dayan, Yam Eitan, Haggai Maron, Soledad Villar·· 14 小时前AI 评分31
Transferable Graph Metanetworks:跨宽度可迁移的图元网络
Transferable Graph Metanetworks
AI 导读
研究者提出 Transferable Graph Metanetworks,通过一组修改让图元网络在不同宽度输入网络间实现性能迁移,在每项任务上显著提升尺寸泛化能力。在 μP 参数化训练的输入网络上表现最强,可稳健泛化至训练宽度的 42 倍。理论部分用无限宽度极限理论证明了 μP 输入下的尺寸泛化保证,并解释了其他参数化下失效的原因。
正文
Abstract:A weight space network (or metanetwork) takes the weights of another neural network as input and predicts properties of it. Most prior work trains such models on input networks of one or a few fixed sizes and evaluates them in-distribution. The few attempts at out-of-distribution size generalization remain limited in scope and have achieved only modest success. Consequently, the potential efficiency gains of training on small networks and evaluating on much larger ones remain largely unrealized. We propose Transferable Graph Metanetworks, which extend the graph metanetwork paradigm with a set of modifications that make performance transferable across input networks of different widths. The modifications follow two principles: invariance to the ways in which networks of different widths represent the same function, and continuity, such that weights representing similar functions receive similar predictions. We further study whether size generalization is possible for input networks trained independently from random initialization. Empirically, our modifications significantly improve size generalization on every task we consider. Performance is strongest on input networks trained under the maximal-update parameterization ($\mu$P), where it remains robust up to $42\times$ the training width. Theoretically, we explain these observations with infinite-width limit theory: we prove size-generalization guarantees for our model on $\mu$P-trained inputs, and explain why it can fail under other parameterizations.
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00420 [stat.ML] |
| (or arXiv:2610.00420v1 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00420 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Soledad Villar [view email]
[v1]
Wed, 30 Sep 2026 15:02:57 UTC (599 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org