arXiv:cs.LG· Samuel Gruffaz, Muhammad Fawad, Jaakko Nevalainen·· 4 小时前AI 评分32
极值二分类:假阴性率极低约束下的极值理论方法
Extreme Binary Classification: Extreme Value Theory for Extreme Constraint on False Negative
AI 导读
研究者提出"极值二分类"问题,目标是学习假阴性率 α 受 ε_N₁=o(1/N₁) 约束的分类器,其中 N₁ 为训练集正样本数。为此提出基于极值理论保证的阈值自适应方法,以及基于样本极大值置换检验的特征选择流程。在四个不同规模真实数据集上,该方法优于现有 SOTA 方法,并在癌症筛查数据集上展示了可解释性。
正文
Abstract:While binary classification is one of the most extensively studied problems in machine learning,
the regime in which the goal is to learn a classifier with an almost zero false negative rate remains largely unexplored.
In this paper, we introduce the Extreme Binary Classification problem, where the objective is to learn a classifier whose false negative rate $\alpha$ is constrained by $\epsilon_{N_1}=o_{N_1\to\infty}(1/N_1)$, with $N_1$ denoting the number of positive examples in the training set.
To address this problem, we propose a threshold adaptation method theoretically grounded in guarantees derived from Extreme Value Theory, together with a feature selection procedure based on a permutation test applied to sample maxima.
Experimental results on four real-world datasets of varying sizes demonstrate that our approach compares favorably with state-of-the-art methods.
In addition, we illustrate its interpretability through an application to a cancer screening dataset.
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG); Applications (stat.AP); Methodology (stat.ME) |
| MSC classes: | 62Cxx |
| Cite as: | arXiv:2610.09984 [stat.ML] |
| (or arXiv:2610.09984v1 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09984 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Samuel Gruffaz [view email]
[v1]
Wed, 7 Oct 2026 12:49:42 UTC (292 KB)
来源:arXiv:cs.LG · arxiv.org