arXiv:cs.LG· Martin Eppert, Krishna Balasubramanian, Subhro Ghosh, Jason Klusowski, Yan Shuo Tan·· 4 小时前AI 评分34
PFN 如何通过上下文学习实现自适应均值估计:一项梯度流分析
Adaptive Mean Estimation by In-Context Learning: A Gradient-Flow Analysis
AI 导读
研究揭示 Prior Fitted Networks(如 TabPFN)在位置估计任务中学习统计自适应的机制:softmax 注意力在标量输入上计算经验累积量生成函数的导数,同时提供区分数据族的特征并在样本均值与中程之间插值。
正文
Abstract:Prior Fitted Networks (PFNs) such as TabPFN now rival established statistical procedures across prediction and estimation tasks. A natural explanation is that PFNs have the property of statistical adaptivity, that is, they perform nearly as well as a method tailored to the true data-generating model for a heterogeneous set of models, while not being told which model the data comes from. We study how such adaptivity is learned in a controlled location-estimation problem. Each task is an unlabeled sample whose family is hidden: Gaussian data call for averaging, with error of order $n^{-1}$, whereas uniform data are best estimated from their extremes, at the faster rate $n^{-2}$. We also provide the example of a symmetric Gaussian mixture, for which a rate of $\sigma^2_n/n$ can be attained. On scalar inputs, softmax attention computes the derivative of the empirical cumulant-generating function. A single primitive therefore both supplies features that distinguish the families and forms estimators interpolating between the sample mean and the mid-range. We combine attention experts through either a softmax mixture of experts or a gated linear unit (GLU), and analyze stagewise gradient flow. With $\widetilde{\Omega}(n^{1+\epsilon})$ pretraining tasks, the learned estimator is asymptotically efficient on Gaussian tasks, within a factor $n^{\epsilon}$ of the minimax rate on uniform tasks, and order-optimal on mixtures in a shrinking-variance regime. These guarantees extend to new locations and longer contexts. A risk decomposition separates expert error, routing error and normalization error, which clarifies the architectural contrast. Softmax gating enforces normalization and exact translation equivariance, whereas the GLU must learn it: its dynamics separate into fast bias removal followed by slow expert selection. End-to-end experiments recover the predicted specialization.
| Subjects: | Machine Learning (cs.LG); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.07804 [cs.LG] |
| (or arXiv:2610.07804v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07804 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Martin Eppert [view email]
[v1]
Tue, 6 Oct 2026 05:57:10 UTC (1,102 KB)
来源:arXiv:cs.LG · arxiv.org