跳到正文
arXiv:cs.LG· Shivam Shrivastava·· 3 小时前AI 评分45

深度表格生成模型在多大训练规模下不再优于平凡基线?一项基于临床与标准数据集规模阶梯的预注册基准测试

Below what training size do deep tabular generators stop beating trivial baselines? A preregistered benchmark on a size ladder of clinical and standard datasets

AI 导读

一项预注册基准在 8 个公开数据集(200 至 20,000 训练行)和 4 个原生小型临床数据集上测试了 7 种生成器,共 2,220 次运行。在 24 个(数据集,深度模型)组合中有 23 个,深度模型在任何训练规模下都未超过最佳平凡基线(差距大于种子噪声),最佳基线在 49 个(数据集,规模)单元中胜出 40 个。

正文

View PDF HTML (experimental)

Abstract:Deep tabular generative models are benchmarked on datasets with tens of thousands of rows; clinical datasets have hundreds. We preregistered and ran a size-ladder benchmark to find where the two regimes diverge: 8 public datasets subsampled from 200 to 20,000 training rows, seven generators (independent marginals, Gaussian copula, SMOTE, unconditional SMOTE, CTGAN, TVAE, TabDDPM) with a fixed 20-trial tuning budget and 5 evaluation seeds, plus 4 natively small clinical datasets at true size, for 2,220 committed runs in total. The primary metric is the AUROC of fixed classifiers trained on synthetic and tested on real data. In 23 of 24 (dataset, deep model) pairs no deep model ever beats the best trivial baseline by more than seed noise, at any training size we measured. The best baseline wins 40 of 49 (dataset, size) cells. Our preregistered prediction that the deep models' ranking would be unstable at small sizes is falsified: mean Kendall tau between adjacent rungs below 5,000 rows is 0.806, above our 0.8 threshold, and stability is highest at the smallest sizes rather than lowest. One caveat bounds all of this: in 81% of cells the gap between the top two methods is smaller than the variation between seeds. Finally, method rankings on natively small clinical datasets agree only moderately with rankings on subsampled large ones (mean tau 0.57 to 0.64), which questions whether a subsampled large dataset can stand in for a small one. All 2,220 result files, the preregistration and its hash, and the code that regenerates every figure and number from those files are public.
Comments: 31 pages, 5 figures. Code, all 2,220 result files and the frozen preregistration: this https URL ; archived at this https URL
Subjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as: arXiv:2610.03500 [cs.LG]
  (or arXiv:2610.03500v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03500

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shivam Shrivastava [view email]
[v1] Fri, 2 Oct 2026 15:58:41 UTC (223 KB)

来源:arXiv:cs.LG · arxiv.org