arXiv:cs.LG· Mikhail Liashkov, Ilyas Varshavskiy, Shuhratjon Khalilbekov, Azizjon Azimi, Bonu Boboeva·· 3 小时前AI 评分44
PaMIR:公共信贷违约数据集开放基准发布 0.4.0
PaMIR: Open Benchmark of Public Credit-Default Datasets
AI 导读
PaMIR 是面向标签稀缺且延迟到达场景的信贷违约预测开放基准,汇集 19 个带二分类违约标签的公共数据集,覆盖九个国家 124 万笔贷款、企业和信用卡账户。所有模型以单一函数形式在重复 i.i.d. 划分和标签延迟流下评分,按标签预算报告 AUC,并配套合成数据测试框架。此次发布为 0.4.0 版本。
正文
Abstract:We release PaMIR (Public Arrival-ordered Measurement for Inference in Risk), an open benchmark for credit-default prediction when labels are scarce and arrive late. The field's reference benchmark studies use eight datasets each, only two or four of them public. PaMIR brings together 19 public datasets with binary default labels -- 1.24M loans, firms and card accounts from nine countries -- rebuilt from pinned source snapshots by one leakage-audited recipe and never redistributed; to our knowledge it is the one of its kind as of today. Every model is a single function, scored under a repeated i.i.d. split and a label-delayed stream in which each application is scored on arrival, with AUC reported by label budget; fleet means are withheld unless every dataset is scored. A synthetic-data harness tests generated training rows without letting a generator see held-out rows. This report describes release 0.4.0 of this living benchmark.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.03259 [cs.LG] |
| (or arXiv:2610.03259v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03259 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shuhratjon Khalilbekov [view email]
[v1]
Fri, 2 Oct 2026 13:04:45 UTC (525 KB)
来源:arXiv:cs.LG · arxiv.org