跳到正文
arXiv:cs.LG· Valentijn Oldenburg, Floris de Kam, Bente Zuijdam, Lieve Eberson, Nicky van Zutphen, Stef de Wildt, Ivo Verhoeven, Cees Snoek·· 4 小时前AI 评分34

流形约束超连接(mHC)用于参数高效微调,冻结主干下动态残差路由

Manifold-Constrained Hyper-Connections for Parameter-Efficient Finetuning

AI 导读

研究者将原本用于预训练的流形约束超连接(mHC)迁移到冻结主干的微调场景,把 Transformer 变为输入依赖的多流残差架构,在每个子层间路由表示。结果显示动态残差路由能完成微调,但作用与预训练不同:保留流混合、仅学习子层如何访问各流,可同时改善损失并减少可训练参数。该工作发表于 NeurIPS 2026 AXIOM 高效深度学习基础 workshop。

正文

View PDF HTML (experimental)

Abstract:Finetuning methods for foundation models usually change weights, prompts, or hidden states, while leaving the residual topology fixed. We ask whether residual topology itself can become a finetuning object. To study this, we adapt manifold-constrained hyper-connections (mHC), recently introduced for pre-training, to frozen-backbone finetuning. mHC turns a Transformer into an input-dependent multi-stream residual architecture, routing representations through multiple streams at every sub-layer. Across mHC variants, we find that dynamic residual routing can finetune Transformers, but that its role differs from pre-training: by preserving stream mixing and learning only how sub-layers access streams, loss is improved and trainable parameters are reduced. Overall, our results identify residual routing as a promising architectural axis for efficient finetuning of foundation models.
Comments: NeurIPS 2026, AXIOM: Foundations of Efficient Deep Learning workshop
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2607.18130 [cs.LG]
  (or arXiv:2607.18130v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2607.18130

arXiv-issued DOI via DataCite

Submission history

From: Valentijn Oldenburg [view email]
[v1] Mon, 20 Jul 2026 16:24:17 UTC (1,159 KB)
[v2] Wed, 7 Oct 2026 08:38:44 UTC (1,128 KB)

来源:arXiv:cs.LG · arxiv.org