跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Adir Dayan, Yam Eitan, Haggai Maron·· 14 小时前AI 评分38

CrossGMN:面向跨架构权重空间变换的图元网络

CrossGMN: Graph Metanetworks for Cross-Architecture Weight-Space Transformations

AI 导读

研究者提出 CrossGMN,一种图元网络,通过对称性保持的跨网络消息传递同时处理源网络与目标网络,实现跨架构权重空间变换。该方法将跨架构算子重构为「已训练源网络 + 目标网络初始化」双输入形式,并证明其在紧集上对连续跨架构算子具有通用性。

正文

View PDF HTML (experimental)

Abstract:Weight-space networks operate directly on parameters of other neural networks, enabling tasks such as predicting model properties, editing trained models, and generating weights. Weight-space symmetries such as neuron permutations make equivariance a key design principle. However, existing equivariant weight-space architectures have primarily been studied for transformations that preserve the network architecture. In contrast, many practical transformations, including model compression and upscaling, map a trained source network into a target network with a different architecture. In this setting, the source and target permutation symmetries act on different parameter spaces, making equivariance less straightforward to formulate. Our key idea for addressing this mismatch is to reformulate cross-architecture operators with two inputs: a trained source network and an initialization of the target network. This lets us define equivariant cross-architecture operators that refine the initialization of the target network using information from the source network, while being invariant to source-network permutations and equivariant to target-network permutations. Based on this formulation, we introduce CrossGMN, a graph metanetwork that jointly processes both networks through symmetry-preserving cross-network message passing. We prove CrossGMN is universal for continuous cross-architecture operators on compact sets under a general-position assumption. We evaluate CrossGMN for model compression, predicting a smaller network's parameters to accelerate subsequent knowledge distillation. Across 2-D and 3-D INRs and image classification with MLPs, CNNs, and Vision Transformers, CrossGMN speeds up distillation by up to 8.89x, transfers across datasets without retraining (3.78x), and a single model can accelerate compression from heterogeneous source architectures into a common target architecture.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.01649 [cs.LG]
  (or arXiv:2610.01649v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01649

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Adir Dayan [view email]
[v1] Thu, 1 Oct 2026 13:13:55 UTC (472 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org