arXiv:cs.LG(机器学习,全量分类)· Soumyadeep Roy·· 14 小时前AI 评分46
IBM Telco 客户流失基准的可信度审计:99% 准确率背后的隐藏代价
The Hidden Costs of 99% Accuracy: A Trustworthiness Audit of the Telco Customer Churn Benchmark
AI 导读
一项对 IBM Telco Customer Churn 基准(n=7,043)的审计发现四类被准确率指标掩盖的可信度问题:预划分 SMOTE 在十种分类器和十五个随机种子下将流失类 F1 虚高 13.1 个百分点(Wilcoxon p<0.05)。
正文
Abstract:Customer churn prediction on the IBM Telco Customer Churn benchmark (n = 7,043) routinely reports test accuracies above 95%, with the most cited published study reporting 99.01%. We audit this benchmark for four trustworthiness failures invisible to the accuracy- and F1-centred reporting that dominates the literature. First, pre-split SMOTE inflates churn-class F1 by 13.1 percentage points across ten classifiers and fifteen seeds (Wilcoxon p < 10^-4 per classifier); the same leaky pipeline ordering paired with class weighting yields no inflation, isolating the effect to SMOTE's geometric construction. We measure the mechanism directly: approximately 36% of synthetic training points are nearest-neighbour interpolations of test-set instances. Second, the TotalCharges field is approximately determined by tenure multiplied by MonthlyCharges (R2 = 0.999); removing it changes accuracy by less than 0.2 percentage points, yet TreeSHAP ranks it ninth in mean absolute attribution - a pattern that materially corrupts SHAP-based interpretation. We propose an R2 > 0.95 pre-modelling diagnostic. Third, in a 15-seed calibration audit, isotonic regression is the strongest default; temperature scaling fails on class-weighted tree ensembles whose predicted-probability distribution is bimodal. Fourth, the cost-optimal decision threshold (under a 50 USD retention offer and 24-month CLV proxy) is approximately 5-10 times lower than the F1-optimal threshold, saving approximately 77,000 USD per 1,000 customers. We replicate F1 and F2 on Iranian Telecom Churn (within domain) and Bank Customer Churn (across domain): F1 generalises; F2 generalises only within telecom. We synthesise these findings into a four-component reporting checklist - pipeline disclosure, redundancy diagnostic, calibration audit, and cost-sensitive thresholds - and release a reproducible implementation.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00118 [cs.LG] |
| (or arXiv:2610.00118v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00118 arXiv-issued DOI via DataCite |
Submission history
From: Soumyadeep Roy [view email]
[v1]
Wed, 9 Sep 2026 12:05:40 UTC (38 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org