跳到正文
arXiv:cs.LG· Pranav Venkata Konda·· 4 小时前AI 评分40

深度归一化注意力的可辨识性与可观测性

The Identifiability and Observability of Deep Normalized Attention

AI 导读

研究证明深度无掩码单头注意力的输入-输出函数可唯一确定有效得分和组合值映射(至偶归一化器诱导的符号),验证了 Henry--Marchetti--Kohn 猜想在实解析情形下成立(含 softmax)。在查询/键同时坍缩时,对常见首个非恒定归一化器阶数 k,第 i 层的接触阶为 2k3^{i-1}-1,并给出精确重数与核维度。

正文

View PDF HTML (experimental)

Abstract:We study which parameters of deep, unmasked, single-head attention are determined by its input--output function. For known positive nonconstant real-analytic normalizers, the function generically determines the effective scores and combined value map up to the signs induced by even normalizers. This proves the real-analytic case of a conjecture of Henry--Marchetti--Kohn, including softmax. We then classify exceptional fibers under explicit normalizer conditions, identifying when collapse makes later scores unobservable, and establish sharp Taylor orders for local identification. Near simultaneous query/key collapse, we compute the complete native Jacobian decay spectrum on separating finite input banks. For common first nonconstant normalizer degree $k$, layer $i$ has contact order $2k3^{i-1}-1$, with exact multiplicities and kernel dimension. High-precision and automatic differentiation calculations illustrate the resulting loss of numerical sensitivity.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.09620 [cs.LG]
  (or arXiv:2610.09620v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09620

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Pranav Venkata Konda [view email]
[v1] Wed, 7 Oct 2026 07:59:10 UTC (96 KB)

来源:arXiv:cs.LG · arxiv.org