跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Saram Abbas, David Thomas, Naeem Soomro, Rishad Shafik, Rakesh Heer, Kabita Adhikari·· 14 小时前AI 评分47

膀胱癌复发预测中的单调约束:修正 XGBoost 学到的“癌症越重风险越低”反转

Fixing a Model That Learned Worse Cancer Means Lower Risk: Monotonic Constraints in Bladder Cancer Recurrence Prediction

AI 导读

针对 UK 多中心 BOXIT 试验中无约束 XGBoost 模型学出“肿瘤分期和原位癌越高、复发风险越低”的反转,研究者提出反事实方向检验与单调约束框架。该模型在 90.2% 的原位癌反事实和 74.3% 的分期反事实中方向反转,而 AUC 差异(ΔAUC 0.005,p=0.47)、校准和 SHAP 幅度均未察觉。

正文

View PDF HTML (experimental)

Abstract:Background and Objective: Clinicians expect recurrence risk to climb with cancer severity. In a UK multicentre trial, an unconstrained XGBoost model learnt that higher tumour stage and carcinoma in situ predicted lower recurrence risk, and discrimination, calibration, and SHAP were all blind to it. We developed a counterfactual testing framework to detect this inversion and a monotonic-constraint framework to remove it without hurting performance. Methods: BOXIT enrolled 472 patients with protocol-mandated cystoscopy across 51 UK sites (2007-2012); 435 had at least two years' follow-up (153 recurrences, 35.2%). We developed a counterfactual direction test and a monotonic-constraint correction, with constraint directions drawn from the EORTC and EAU risk systems, and evaluated both against unconstrained XGBoost and logistic regression on 18 predictors (seven directed) over 50 cross-validation folds. The test worsened each patient on one directed feature at a time to check whether risk fell; SHAP direction and calibration were also assessed. Key Findings and Limitations: Tumour stage and carcinoma in situ were associated with lower recurrence, opposite to medical intuition; the unconstrained model reversed carcinoma in situ counterfactuals in 90.2% of cases and stage in 74.3%. Discrimination ($\Delta$AUC 0.005, p=0.47), calibration, and SHAP magnitude were all blind to the inversion. Monotonic constraints eliminated every violation at no cost to discrimination (0.723 vs 0.718) and outperformed EORTC (p=8.9e-16). Limitations: single trial, internal-external validation only. Conclusions and Clinical Implications: A model that had learned this inversion passed every conventional check. A counterfactual direction test, run as a single refit with pre-specified monotonic constraints, catches this failure at no cost to performance and should be routine before clinical deployment.
Comments: 13 pages, 4 figures, 1 table. Supplementary material (15 pages) provided as ancillary files
Subjects: Machine Learning (cs.LG)
MSC classes: 62P10, 68T05
ACM classes: I.2.6; J.3
Cite as: arXiv:2610.00858 [cs.LG]
  (or arXiv:2610.00858v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00858

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Saram Abbas [view email]
[v1] Thu, 1 Oct 2026 00:21:33 UTC (1,616 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org