跳到正文
arXiv:cs.LG· Kaustubh Sharma, Srijan Tiwari, Ojasva Nema, Parikshit Pareek·· 3 小时前

PFN 中频谱结构的机制证据:TabPFN 等模型内部可读出显式核

Mechanistic Evidence for Spectral Structures in Prior-Data Fitted Networks

AI 导读

研究者在七个 PFN(含四个预训练 TFM 与一个仅用决策树先验训练的模型)的残差流上做线性探针,以 R²≥0.95 恢复上下文的频率,且该结构由单一主方向主导。

正文

View PDF HTML (experimental)

Abstract:Prior-Data Fitted Networks (PFNs) perform approximate Bayesian inference in a single forward pass, and tabular foundation models (TFMs) built on them are now widely used. To understand what networks infer internally, recent mechanistic studies of TFMs locate where predictions form, but treat these models as tabular predictors rather than as PFNs. It therefore remains unknown whether PFNs represent the spectral content of their context, the quantity that specifies a stationary kernel, and whether this content can be read out as an explicit kernel. We answer both questions. First, across seven PFNs, including four pretrained TFMs and a model trained only on a decision-tree prior, a linear probe on the residual stream recovers the frequency of the context with $R^2 \geq 0.95$. This structure is led by a single principal direction. Second, activation and subspace patching show that the network uses the structure through a low-dimensional subspace, where a few spectral directions move predictions far more than random ones. This holds even for the decision-tree model, so a spectral training prior is not required. On real datasets with up to 499 features, 64 of the 192 directions of TabPFN, chosen without labels, carry 85 to 95\% of the causal effect of the context in all but one pair. Third, we introduce a Filter Bank Decoder that turns frozen PFN representations into an explicit stationary kernel through Bochner's theorem. Without any test-time optimization, the decoded kernel supports Gaussian process regression competitive with deep kernel learning and random Fourier features at about $200\times$ lower cost. PFN latents therefore hold spectral structure that is causally used and recoverable as a portable kernel.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2601.21731 [cs.LG]
  (or arXiv:2601.21731v3 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2601.21731

arXiv-issued DOI via DataCite

Submission history

From: Parikshit Pareek [view email]
[v1] Thu, 29 Jan 2026 13:51:26 UTC (1,958 KB)
[v2] Wed, 13 May 2026 02:43:09 UTC (1,408 KB)
[v3] Thu, 8 Oct 2026 13:12:23 UTC (1,376 KB)

来源:arXiv:cs.LG · arxiv.org