arXiv:cs.LG(机器学习,全量分类)· Kazuma Sawaya·· 15 小时前AI 评分36
深度神经网络特征选择的 FDR 控制理论证明:Deep MLPs 及更广架构
Provable FDR Control for Deep Feature Selection: Deep MLPs and Beyond
AI 导读
研究者提出一种基于深度神经网络的特征选择框架,可近似控制错误发现率(FDR),适用于首层全连接、后续可为任意宽度深度 MLP、卷积、循环网络、注意力、残差连接和 dropout 的架构。该工作首次在此类通用深度学习设定下给出 FDR 控制的理论保证,基于多指标数据生成模型,证明梯度特征重要性向量各坐标满足边际正态近似。论文已被 AISTATS 2026 接收。
正文
Abstract:We develop a flexible feature selection framework based on deep neural networks that approximately controls the false discovery rate (FDR), a measure of Type-I error. The method applies to architectures whose first layer is fully connected. From the second layer onward, it accommodates multilayer perceptrons (MLPs) of arbitrary width and depth, convolutional and recurrent networks, attention mechanisms, residual connections, and dropout. The procedure also accommodates stochastic gradient descent with data-independent initializations and learning rates. To the best of our knowledge, this is the first work to provide a theoretical guarantee of FDR control for feature selection within such a general deep learning setting.
Our analysis is built upon a multi-index data-generating model and an asymptotic regime in which the feature dimension $n$ diverges faster than the latent dimension $q^{*}$, while the sample size, the number of training iterations, the network depth, and hidden layer widths are left unrestricted. Under this setting, we show that each coordinate of the gradient-based feature-importance vector admits a marginal normal approximation, thereby supporting the validity of asymptotic FDR control. As a theoretical limitation, we assume $\mathbf{B}$-right orthogonal invariance of the design matrix, and we discuss broader generalizations. We also present numerical experiments that underscore the theoretical findings.
| Comments: | Accepted to AISTATS 2026 |
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST) |
| Cite as: | arXiv:2512.04696 [stat.ML] |
| (or arXiv:2512.04696v3 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2512.04696 arXiv-issued DOI via DataCite |
Submission history
From: Kazuma Sawaya [view email]
[v1]
Thu, 4 Dec 2025 11:46:06 UTC (490 KB)
[v2]
Mon, 9 Feb 2026 10:44:45 UTC (491 KB)
[v3]
Thu, 1 Oct 2026 01:15:26 UTC (780 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org