跳到正文
arXiv:cs.AI· Minjia Mao, Shi Chen, Bowen Yin, Xiao Fang·· 6 小时前AI 评分43

大语言模型中的大规模激活门控通道(MAGC)

Massive Activation Gating Channel in Large Language Models

AI 导读

研究发现,LLM 中大规模激活的出现由输入嵌入到 spike FFN 的单一通道控制,该通道位置对特定 LLM 固定,被命名为大规模激活门控通道(MAGC)。当 MAGC 值足够大(或足够小,视 LLM 而定)时,spike FFN 输出呈现大规模激活;在四个模型家族、不同规模的六个 LLM 上验证了 MAGC 的存在与作用。

正文

View PDF HTML (experimental)

Abstract:Massive activations, a phenomenon in which a small number of hidden channels exhibit exceptionally large magnitudes, are pervasive in large language models (LLMs). However, the mechanism by which a token develops massive activations as it propagates through a pretrained LLM remains poorly understood. In this paper, we find that the emergence of massive activations is controlled by a single channel in the input embedding to a spike feed-forward network (FFN). The position of this channel is fixed for a particular LLM. We name this channel the massive activation gating channel (MAGC). When the value of the MAGC is sufficiently large (or small, depending on the LLM), the output of the spike FFN exhibits massive activations. Examining six LLMs across four model families and different model sizes, we verify the existence and effect of MAGC. We further provide a theoretical explanation of the mechanism by which MAGC induces massive activations. When the value of MAGC is sufficiently large (or small), the output of a spike FFN asymptotically reduces to a quadratic form that mixes a few columns of the down-projection matrix of the FFN. Since these columns exhibit the shape of massive activations, the output therefore exhibits massive activations.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07661 [cs.AI]
  (or arXiv:2610.07661v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07661

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Minjia Mao [view email]
[v1] Tue, 6 Oct 2026 02:55:48 UTC (4,847 KB)

来源:arXiv:cs.AI · arxiv.org