arXiv:cs.LG· Walid Bendada, Guillaume Salha-Galvan·· 4 小时前AI 评分35
两级 Softmax 采样修正:S-2LS 与 SD-2LS 消除规模失衡与分散度偏差
Two-Level Softmax Sampling Done Right: Correcting Bias from Size Imbalance and Dispersion
AI 导读
针对两级 Softmax(2LS)采样因忽略簇规模失衡和簇内相似度分散度而引入的系统性偏差,研究者提出 S-2LS 和 SD-2LS 两种采样方法,可证明地给出更优的 softmax 近似且计算开销可忽略。在五个大规模数据集上的实验验证了改进的采样性质,作者建议未来工作统一用其替代标准 2LS。该成果已被 NeurIPS 2026 接收。
正文
Abstract:Sampling from a softmax distribution is a fundamental operation in machine learning, but its linear complexity in the number of items makes exact sampling impractical at scale. Two-level softmax (2LS) sampling is a popular alternative enabling sublinear-time sampling. Assuming items are partitioned into clusters, 2LS first samples a cluster and then an item within it. In this paper, we show that, despite its advantages, 2LS introduces systematic and undesirable sampling biases, which arise from misweighting clusters by ignoring both cluster size imbalance and intra-cluster similarity dispersion. We propose two sampling methods, Size-Corrected 2LS (S-2LS) and Size- and Dispersion-Corrected 2LS (SD-2LS), which correct these biases and provide provably better softmax approximations with negligible to non-existent computational overhead. In-depth experiments on five large-scale datasets validate the improved sampling properties of our methods. We recommend their consistent use in place of standard 2LS in future work.
| Comments: | NeurIPS 2026 |
| Subjects: | Machine Learning (cs.LG); Information Retrieval (cs.IR); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.10483 [cs.LG] |
| (or arXiv:2610.10483v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10483 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Guillaume Salha-Galvan [view email]
[v1]
Wed, 7 Oct 2026 17:38:35 UTC (1,202 KB)
来源:arXiv:cs.LG · arxiv.org