arXiv:cs.LG· Cagdas Pullu, Mahmut Emir Arslan, Bugra Balkac, Aylin Ondersev Balta, Cihangir Celal Palaci, Fikri Cem Yilmaz, Altan Cakir·· 5 小时前AI 评分34
面向大规模电商运营数据概念漂移的分布式联合分布检验
Beyond Marginal Monitoring: Distributed Joint-Distribution Testing for Data Concept Drift in Large Scale E-Commerce Operations
AI 导读
研究在 1.375 亿行的 Trendyol 集合排序特征表上评估五种多列双样本漂移检测方法,发现基于 Apache Spark 的分布式最大均值差异(MMD)配合随机傅里叶特征扩展最稳健。
正文
Abstract:Concept drift threatens production machine learning, yet the empirical behavior of multivariate two-sample drift detectors at scale remains under-characterized. Existing benchmarks rarely address the hundreds of millions of rows and high-cardinality features typical of industrial-operational datasets. We evaluate five multi-column two-sample tests (marginal, projection-based, and kernel embedding methods) across three complementary environments: the Harvard Dataverse, a validated Failing Loudly reproduction (mean absolute error between 0.030 and 0.053), and a novel synthetic-injection benchmark on the 137.5-million-row Trendyol collection-ranking feature table. Testing four drift types across two severity-scope regimes, we demonstrate that distributed Maximum Mean Discrepancy with Random Fourier Features on Apache Spark scales robustly. Averaged over the four drift types in the strong regime and under a calibrated threshold, it achieves a Pearson correlation of r = 0.940 with expected drift magnitude, an 80.4% true positive rate, and a 3.2% false positive rate. Conversely, the per-dimension Kolmogorov-Smirnov test failed due to statistic saturation from ID-like columns under asymmetric sampling, establishing a critical constraint for large-scale sampling design. At weak configurations (realized-flip fractions of at most 0.57%), detectors struggled to reliably discriminate, highlighting the need for future intensity-grid power analyses to distinguish fundamental sensitivity bounds from scalable threshold shifts.
| Subjects: | Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG); Methodology (stat.ME); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.08132 [cs.DC] |
| (or arXiv:2610.08132v1 [cs.DC] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08132 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Altan Cakir [view email]
[v1]
Tue, 6 Oct 2026 10:47:26 UTC (3,664 KB)
来源:arXiv:cs.LG · arxiv.org