arXiv:cs.LG· Peng Zeng, Hanwen Huang·· 7 小时前AI 评分27
稀疏支持向量机的高维统计推断
High-Dimensional Statistical Inference for Sparse Support Vector Machines
AI 导读
研究提出一种针对稀疏支持向量机的高维推断框架,在样本量与特征数成比例增长时,将 L1 惩罚 SVM 表示为线性规划并通过其对偶变量识别 hinge 损失的次梯度,从而得到坐标渐近高斯的去偏估计量。
正文
Abstract:Using a replica-symmetric high-dimensional characterization, we develop an inferential framework for sparse support vector machines when the sample size and number of features grow proportionally. The main challenge is the nonsmooth hinge loss, which prevents direct application of debiasing arguments developed for smooth classification losses. We overcome this difficulty by representing the $L_1$-penalized support vector machine (SVM) as a linear program and identifying the hinge-loss subgradient through its dual variables. This yields a computationally accessible debiased estimator whose coordinates are asymptotically Gaussian under the proportional asymptotic regime. The resulting distributional characterization provides confidence intervals and hypothesis tests for individual features and enables false-discovery-rate-controlled variable selection. Extensive simulations examine calibration, power, and variable-selection performance under a range of covariance structures, including strongly correlated designs. An analysis of high-dimensional breast cancer gene-expression data illustrates how the proposed inference can distinguish statistically significant features from variables selected by the original sparse SVM.
| Comments: | 7 figures |
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG); Methodology (stat.ME) |
| MSC classes: | 62H30 |
| Cite as: | arXiv:2610.08345 [stat.ML] |
| (or arXiv:2610.08345v1 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08345 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Peng Zeng [view email]
[v1]
Tue, 6 Oct 2026 13:40:12 UTC (120 KB)
来源:arXiv:cs.LG · arxiv.org