arXiv:cs.LG· Zhihan Huang, Ziang Niu·· 4 小时前AI 评分36
基于 kernelized Stein 差异的高计算效率拟合优度检验
Computationally efficient goodness-of-fit tests through kernelized Stein discrepancy
AI 导读
研究者提出一种基于 kernelized Stein 差异(SKSD)的半参数拟合优度检验,无需重新拟合模型或从模型采样即可构造 level-α 检验。该检验在经典正态性检验到难解似然模型的模拟中,以低数个数量级的计算成本取得相当或更优的检验功效,并用于评估肺腺癌反向蛋白阵列数据的蛋白信号网络模型。研究还证明 SKSD 检验可视为指数倾斜模型下的非参数得分检验。
正文
Abstract:Models with intractable normalizing constants are widely used in statistics and machine learning. Assessing the adequacy of such models poses significant challenges: obtaining samples from the fitted model often requires sophisticated sampling algorithms. Moreover, model fitting sometimes requires iterative numerical optimization, making bootstrap procedures that require repeated refitting computationally expensive. In this paper, we leverage the kernel-based testing framework to develop a general semiparametric goodness-of-fit test based on the kernelized Stein discrepancy. We establish the consistency and the asymptotic null distribution of the test statistic under general nuisance estimation. To produce a level-$\alpha$ test, we propose a novel influence-adjusted wild bootstrap that requires neither refitting the model nor sampling from it. We prove the consistency of the proposed bootstrap test procedure under the null and the alternative, and characterize its limiting power under contiguous local alternatives. Across simulations ranging from classical normality testing to models with intractable likelihoods, the proposed test delivers competitive or superior power at a computational cost orders of magnitude lower than that of existing approaches. We illustrate the method by assessing the adequacy of a protein signaling network model for reverse-phase protein array data from lung adenocarcinoma tumors. As a complementary insight, we show that the SKSD test can be regarded as a nonparametric score test under exponentially tilted models, connecting score-based and distance-based goodness-of-fit testing.
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME) |
| Cite as: | arXiv:2512.20007 [stat.ML] |
| (or arXiv:2512.20007v3 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2512.20007 arXiv-issued DOI via DataCite |
Submission history
From: Zhihan Huang [view email]
[v1]
Tue, 23 Dec 2025 03:05:26 UTC (129 KB)
[v2]
Fri, 20 Feb 2026 22:18:17 UTC (121 KB)
[v3]
Tue, 6 Oct 2026 06:17:23 UTC (1,167 KB)
来源:arXiv:cs.LG · arxiv.org