跳到正文
arXiv:cs.LG· Alina Sudakov, Guy Bar-Shalom, Fabrizio Frasca, Haggai Maron·· 3 小时前AI 评分40

MATCHA:用显式多对多层映射学习跨模型激活对齐

Learning Cross-Model Activation Alignments with Explicit Many-to-Many Layer Maps

AI 导读

研究者提出 MATCHA,通过从提示词中联合学习层映射与层共享特征映射,实现源模型到目标模型的激活对齐,其中层映射可输出并检查显式的 target-by-source 矩阵。

正文

View PDF HTML (experimental)

Abstract:LLMs are released at a rapid pace, raising a natural question: how do two independently trained models relate, both in which layers correspond and in how features transform between them? We study this by learning an activation alignment, a map from a source model's layerwise activations to a target's. Our method, MATCHA, factors this map into a layer map, whose output is an explicit target-by-source matrix that can be extracted and inspected, and a layer-shared feature map between hidden spaces. Most of prior work fixes the layer correspondence in advance, pairing layers at roughly the same relative depth; in contrast, we learn both factors jointly from prompts. Across 42 pairs of seven models spanning three different families, MATCHA reconstructs the target's activations more faithfully and improves retrieval-based metrics substantially, w.r.t. previous approaches. The recovered maps are broadly monotone in depth but, in contrast with most previous approaches, are consistently many-to-many: each target layer draws on a band of source layers. Our alignments also enable transfer of activation-space interventions, allowing steering vectors and probes developed for one model to transfer to another.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.09058 [cs.LG]
  (or arXiv:2610.09058v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09058

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Alina Sudakov [view email]
[v1] Tue, 6 Oct 2026 20:08:41 UTC (1,514 KB)

来源:arXiv:cs.LG · arxiv.org