跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Luca Zhou, Emanuele Rodol\`a·· 15 小时前AI 评分54

arXiv 论文研究模型合并对涌现能力的影响

On Emergent Capabilities and Model Merging

AI 导读

arXiv 论文(arXiv:2609.24504)研究模型合并(model merging)对涌现能力的影响,在激活 oracle 和涌现失齐模型两个测试平台、三个模型族上得出三重结论。

正文

View PDF HTML (experimental)

Abstract:Fine-tuned checkpoints and adapters now fill public repositories, and the most common operation applied to these artifacts is model merging: arithmetic on their weights that assembles capabilities cheaply. We ask what this operation does to emergent capabilities: behaviors an artifact carries that were never an explicit training target. Studying two independent testbeds (activation oracles and emergent-misaligned models) across three model families, we find that the answer is threefold. First, merging preserves an emergent capability that both parents carry: merging two misaligned checkpoints retains most of their broad misalignment across the whole mixing range. Second, merging cannot create an emergent capability that is superadditive in its parents: no weighted merge of two single-task oracles reaches the jointly-trained oracle's auditing ability. Third, when only one parent carries the capability, merging dilutes it faster than the trained capability that accompanies it: the gap is significant in most settings. In short, emergent behaviors of an artifact do not compose the way its trained capability does.
Comments: main paper has 8 pages, 5 figures, and 4 tables
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.24504 [cs.LG]
  (or arXiv:2609.24504v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.24504

arXiv-issued DOI via DataCite

Submission history

From: Luca Zhou [view email]
[v1] Mon, 21 Sep 2026 12:44:35 UTC (212 KB)
[v2] Thu, 1 Oct 2026 13:56:04 UTC (212 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org