arXiv:cs.LG· Junghwan Park, Woojin Cho·· 2 天前AI 评分44
Platonic Task Arithmetic:跨异构模型的任务向量迁移与组合
Platonic Task Arithmetic
AI 导读
研究者提出“柏拉图任务向量”假设,认为不同模型针对同一任务的参数更新是同一模型无关对象的投影,并引入形状与架构、嵌入维度无关的 Universal Task Descriptors 来记录任务功能效果。该描述符可通过最小二乘求解折叠进目标模型最后一层,或用低秩适配器联合拟合组合,无需逐图标签。在六个模型族、八项分类任务及音频-文本设置中,跨模型迁移保留了目标自身描述符 74-80% 的增益。
正文
Abstract:Models specialized for the same task converge to similar behavior, yet the parameter updates that produce it share no common coordinate system, so weight-space task arithmetic stays confined to a single model and cannot cross architectures without a structural correspondence. Drawing on Plato's allegory of the cave, we hypothesize that these model-specific updates are shadows of one shared, model-agnostic object, which we call the platonic task vector. To make it operational for models that pair an image or audio encoder with a text encoder, we introduce Universal Task Descriptors: matrices whose shape is independent of architecture and embedding dimension, which record a task's functional effect and support addition and negation as matrix operations. Transferring a descriptor into a target means editing the target until it reproduces the descriptor on the task's unlabeled probe images and class-name prompts, requiring no per-image labels. We realize this edit in two ways. First, the descriptor factorizes into a shift field on image embeddings, so a single least-squares solve yields a linear operator that folds into the target's last layer as a weight edit; by linearity, a bank of such operators admits any composition at any strength as a signed sum. Second, a low-rank adapter trained on the same objective reaches every layer and fits compositions jointly, at the cost of one optimization per edit. Heterogeneous models share this object only partially, with a model-specific residual comparable in norm to the shared component, yet cross-model transfer still retains 74-80 percent of the gain of the target's own descriptors. Experiments across six model families, eight classification tasks, and an audio-text setting show that task knowledge transfers and composes across heterogeneous models under both realizations.
| Comments: | NeurIPS2026 |
| Subjects: | Machine Learning (cs.LG); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.00929 [cs.LG] |
| (or arXiv:2610.00929v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00929 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Junghwan Park [view email]
[v1]
Thu, 1 Oct 2026 02:05:13 UTC (299 KB)
来源:arXiv:cs.LG · arxiv.org