arXiv:cs.LG· Pranav Venkata Konda·· 4 小时前AI 评分40
深度归一化注意力的可辨识性与可观测性
The Identifiability and Observability of Deep Normalized Attention
AI 导读
研究证明深度无掩码单头注意力的输入-输出函数可唯一确定有效得分和组合值映射(至偶归一化器诱导的符号),验证了 Henry--Marchetti--Kohn 猜想在实解析情形下成立(含 softmax)。在查询/键同时坍缩时,对常见首个非恒定归一化器阶数 k,第 i 层的接触阶为 2k3^{i-1}-1,并给出精确重数与核维度。
正文
Abstract:We study which parameters of deep, unmasked, single-head attention are determined by its input--output function. For known positive nonconstant real-analytic normalizers, the function generically determines the effective scores and combined value map up to the signs induced by even normalizers. This proves the real-analytic case of a conjecture of Henry--Marchetti--Kohn, including softmax. We then classify exceptional fibers under explicit normalizer conditions, identifying when collapse makes later scores unobservable, and establish sharp Taylor orders for local identification. Near simultaneous query/key collapse, we compute the complete native Jacobian decay spectrum on separating finite input banks. For common first nonconstant normalizer degree $k$, layer $i$ has contact order $2k3^{i-1}-1$, with exact multiplicities and kernel dimension. High-precision and automatic differentiation calculations illustrate the resulting loss of numerical sensitivity.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09620 [cs.LG] |
| (or arXiv:2610.09620v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09620 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pranav Venkata Konda [view email]
[v1]
Wed, 7 Oct 2026 07:59:10 UTC (96 KB)
来源:arXiv:cs.LG · arxiv.org