arXiv:cs.LG· Antonij Mijoski, Marko Karbevski·· 7 小时前AI 评分37
MLP 能否精确吸收自身的跳跃连接?
Can an MLP Absorb Its Own Skip Connection Exactly?
AI 导读
一项研究追问:给定残差块 x -> x + MLP(x),同宽度的无残差 MLP 能否计算相同函数?答案是在任意深度下无条件为否,对当前前沿语言模型使用的所有激活函数(ReLU^2、ReGLU、SwiGLU、GeGLU)均成立。
正文
Abstract:The benefits usually attributed to skip connections are optimization-theoretic: a smoother loss landscape and better gradient propagation. We ask a representational question instead: given a residual block x -> x + MLP(x), does a residual-free MLP of the same width compute the same function? The answer is no, unconditionally and at every depth, for every activation used in current frontier language models (ReLU^2, ReGLU, SwiGLU, GeGLU). For ungated ReLU and GELU absorption is possible, but only on a set of weights of measure zero. The two families are therefore generically disjoint: removing a skip connection and retraining cannot recover exactly the same function at equal width.
| Comments: | Accepted at the NeurReps Workshop @ NeurIPS 2026 (Extended Abstract Track), nominated for oral presentation. Extended version. OpenReview: this https URL |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2604.23705 [cs.LG] |
| (or arXiv:2604.23705v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2604.23705 arXiv-issued DOI via DataCite |
Submission history
From: Marko Karbevski [view email]
[v1]
Sun, 26 Apr 2026 13:37:27 UTC (20 KB)
[v2]
Mon, 5 Oct 2026 18:52:49 UTC (18 KB)
来源:arXiv:cs.LG · arxiv.org