arXiv:cs.LG· Hong Ha Le, Jackie Lok, Atsushi Nitanda, Yan Shuo Tan·· 4 小时前AI 评分30
软最大注意力如何学习上下文中的决策树桩阈值
Learning Decision-Stump Thresholds in Context: Dynamics of Softmax Attention
AI 导读
研究揭示了两参数softmax-attention模型通过基于梯度的预训练学习决策阈值统计规则的动力学机制。在大分辨率初始化下,对m个任务进行恒定步长梯度下降产生冻结估计器,误差为Õ((m∧n)⁻¹+N⁻¹)。该机制源于协调参数发散:群体训练先校准相对标签与特征分数,再以t^{1/4}增加注意力尺度,使群体阈值误差为O(t^{-1/4})。
正文
Abstract:Estimating a decision threshold requires locating observations near an unknown boundary. We study how gradient-based pretraining learns this statistical rule in a two-parameter softmax-attention model with a fixed feature and inequality direction. Pretraining uses labeled contexts and their true thresholds; a fresh threshold must be inferred from context alone. Under a large-resolution initialization, constant-step gradient descent on $m$ tasks with $n$ examples each produces a frozen estimator with error $\widetilde O((m\wedge n)^{-1}+N^{-1})$ for each fixed interior threshold and every fresh-context size $N$. The two terms separate finite-pretraining accuracy from fresh-context localization. The mechanism is coordinated parameter divergence: population training calibrates the relative label and feature scores, then increases the attention scale as $t^{1/4}$, giving population threshold error $O(t^{-1/4})$. To transfer this mechanism to a fixed finite corpus, we control gradient errors relative to the shrinking directions of progress at successive parameter scales. This certifies a growing training interval without requiring long-time tracking of the population trajectory. We also identify the boundary limitation of the one-head model and explain statistically what a reflected symmetrization could achieve.
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07074 [stat.ML] |
| (or arXiv:2610.07074v1 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07074 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hong Ha Le [view email]
[v1]
Mon, 5 Oct 2026 09:00:28 UTC (233 KB)
来源:arXiv:cs.LG · arxiv.org