arXiv:cs.LG(机器学习,全量分类)· Atsutoshi Kumagai, Tomoharu Iwata, Taishi Nishiyama, Hiroshi Takahashi, Kazuki Adachi, Yasuhiro Fujiwara·· 14 小时前AI 评分27
从正例-无标签数据中最大化部分 AUC
Partial AUC Maximization from Positive-unlabeled Data
AI 导读
针对缺乏标注负样本时难以最大化部分 AUC(pAUC)的问题,研究者提出一种仅用正例和无标签(PU)数据训练分类器的方法。该方法在经验风险最小化框架下证明 pAUC 及其依赖 FPR 的阈值可仅用正例密度和边缘密度表示,并据此推导出 PU 数据的经验估计量,通过最大化平滑后的经验 pAUC 估计量训练分类器。在十个真实数据集上的实验验证了该方法的有效性。
正文
Abstract:The partial area under the receiver operating characteristic curve (pAUC) is an important performance metric for binary classification that summarizes true positive rates within a specific range of false positive rates (FPRs). Classifiers that achieve high pAUC need to be obtained in many real-world applications such as cybersecurity, medical care, and advertising. Although many methods for maximizing the pAUC have been proposed, they typically require both labeled positive and negative data for training. However, in practice, labeled negative data are often difficult to collect due to privacy concerns or the need for high expertise to annotate them. In this paper, we propose a method for maximizing the pAUC from positive and unlabeled (PU) data without negative data. Within an empirical risk minimization framework, we show that the pAUC, including its FPR-dependent thresholds, can be represented using only the positive and marginal densities, and derive an empirical estimator from PU data. A classifier is then trained by maximizing the derived smoothed empirical pAUC estimator. We experimentally demonstrate the effectiveness of the proposed method with ten real-world datasets.
| Comments: | 26 pages |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.00284 [cs.LG] |
| (or arXiv:2610.00284v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00284 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Atsutoshi Kumagai [view email]
[v1]
Fri, 25 Sep 2026 03:43:39 UTC (349 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org