arXiv:cs.LG· Yanxuan Yu, Dong Liu, Shu Wang, Wenxiao Zhao, Eric Jiang, Chang Liu, Jinxi Yu, Hui Pan, Renata Borovica-Gajic, Ben Lengerich·· 4 小时前AI 评分33
RUBRIC:面向不平衡分类的现实性-效用平衡排序框架
RUBRIC: Realism--Utility Balanced Ranking for Imbalanced Classification
AI 导读
RUBRIC 是一个与生成器无关的合成样本过滤框架,将合成样本选择形式化为重质不重量的优化问题,通过现实性与效用的权衡对候选样本排序。现实性由神经密度比判别器估计,效用用基于间隔的凹评分函数 g_-(t)=-log(1+e^{-t/τ}) 衡量其靠近决策边界的程度。
正文
Abstract:Class imbalance poses a fundamental challenge in risk-sensitive applications such as fraud detection and medical diagnosis, where minority-class samples are scarce yet critical for accurate classification. Existing oversampling methods generate synthetic samples to rebalance class distributions; however, they often produce large numbers of low-quality candidates that distort decision boundaries or introduce artifacts, leading to overfitting and degraded generalization.
In this work, we introduce \textbf{RUBRIC}, a generator-agnostic filtering framework that formulates synthetic sample selection as a quality-over-quantity optimization problem. RUBRIC ranks candidates using a realism-utility trade-off: realism is estimated via a neural density-ratio discriminator from each candidate's resemblance to real minority samples, while utility captures proximity to the decision boundary through a concave, margin-based scoring function $g_-(t)=-\log(1+e^{-t/\tau})$. The discriminator uses the same architecture and training protocol on every benchmark and is fit independently to that dataset's real minority class versus its synthetic pool. We show that, under mild regularity conditions, the proposed filtering framework monotonically tightens the generalization bound for margin-based classifiers by jointly reducing distribution shift and suppressing near-negative tail contributions.
Through extensive experiments on standard public imbalanced-classification benchmarks, we demonstrate that RUBRIC boosts minority-class recall while preserving overall discriminative ability across multiple data generators. Sensitivity analyses in $\lambda$ and the selection budget $K$ further characterize performance trade-offs oriented toward ranking quality.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.09816 [cs.LG] |
| (or arXiv:2607.09816v4 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2607.09816 arXiv-issued DOI via DataCite |
Submission history
From: Dong Liu [view email]
[v1]
Fri, 10 Jul 2026 06:37:12 UTC (6,452 KB)
[v2]
Tue, 21 Jul 2026 10:33:53 UTC (6,452 KB)
[v3]
Fri, 14 Aug 2026 02:29:45 UTC (6,470 KB)
[v4]
Wed, 7 Oct 2026 03:19:25 UTC (7,288 KB)
来源:arXiv:cs.LG · arxiv.org