跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Bin Chen, Yumeng Xue, Patrick Paetzold, Yunhai Wang, Oliver Deussen·· 14 小时前AI 评分40

ibUMAP:面向 UMAP 优化的连贯可扩展场评估方法

ibUMAP: Coherent and Scalable Field Evaluation for UMAP Optimization

AI 导读

ibUMAP 提出一种基于场的 UMAP 优化替代方案,从共享嵌入快照同步计算吸引与排斥力,其度加权排斥场由三个标量矩表示,并用基于插值的 FFT 方案在 CPU 和 GPU 上高效评估。

正文

View PDF HTML (experimental)

Abstract:UMAP achieves scalable layout optimization through stochastic negative sampling. However, this stochasticity can lead to unstable embeddings across reruns and downstream reuse, as the estimated repulsive forces depend on the ordering of sampling events. We present ibUMAP, a coherent field-based alternative that evaluates attraction and repulsion from a shared embedding snapshot and applies them synchronously. Its degree-weighted repulsive field is motivated by the conditional expectation of negative sampling for a fixed embedding and represented by three scalar moments, which are evaluated efficiently on CPUs and GPUs using an interpolation-based FFT scheme. This formulation avoids explicit all-pairs computations while inducing optimization dynamics that differ from those of standard online UMAP. Controlled experiments show that synchrony and kernel capping alter the local-global fidelity trade-off, whereas FFT evaluation produces small average changes in final quality. End-to-end benchmarks show median speedups of 3.29x unseeded and 5.79x seeded over umap-learn on CPU, and 1.44x over cuML on million-scale datasets under unseeded GPU execution. These gains accompany greater run-to-run stability and measurable fidelity trade-offs.
Comments: 35 pages, 11 figures, 18 tables. Under review at ICLR 2027. Code: this https URL
Subjects: Machine Learning (cs.LG); Human-Computer Interaction (cs.HC)
Cite as: arXiv:2610.01445 [cs.LG]
  (or arXiv:2610.01445v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01445

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Bin Chen [view email]
[v1] Thu, 1 Oct 2026 10:41:55 UTC (2,387 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org