跳到正文
arXiv:cs.LG· Antonij Mijoski, Marko Karbevski·· 7 小时前AI 评分37

MLP 能否精确吸收自身的跳跃连接?

Can an MLP Absorb Its Own Skip Connection Exactly?

AI 导读

一项研究追问:给定残差块 x -> x + MLP(x),同宽度的无残差 MLP 能否计算相同函数?答案是在任意深度下无条件为否,对当前前沿语言模型使用的所有激活函数(ReLU^2、ReGLU、SwiGLU、GeGLU)均成立。

正文

View PDF HTML (experimental)

Abstract:The benefits usually attributed to skip connections are optimization-theoretic: a smoother loss landscape and better gradient propagation. We ask a representational question instead: given a residual block x -> x + MLP(x), does a residual-free MLP of the same width compute the same function? The answer is no, unconditionally and at every depth, for every activation used in current frontier language models (ReLU^2, ReGLU, SwiGLU, GeGLU). For ungated ReLU and GELU absorption is possible, but only on a set of weights of measure zero. The two families are therefore generically disjoint: removing a skip connection and retraining cannot recover exactly the same function at equal width.
Comments: Accepted at the NeurReps Workshop @ NeurIPS 2026 (Extended Abstract Track), nominated for oral presentation. Extended version. OpenReview: this https URL
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2604.23705 [cs.LG]
  (or arXiv:2604.23705v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2604.23705

arXiv-issued DOI via DataCite

Submission history

From: Marko Karbevski [view email]
[v1] Sun, 26 Apr 2026 13:37:27 UTC (20 KB)
[v2] Mon, 5 Oct 2026 18:52:49 UTC (18 KB)

来源:arXiv:cs.LG · arxiv.org