跳到正文
arXiv:cs.LG· Liang You, Hengyu Shi, Dongwen Ou·· 2 天前AI 评分39

SimplexUQ:面向单纯形值预测的共形不确定性评估框架与基准

SimplexUQ: An Evaluation Framework and Benchmark for Conformal Uncertainty on Simplex-Valued Predictions

AI 导读

SimplexUQ 是首个针对单纯形值预测中共形覆盖分配问题的基准与可复现协议,其任务套件 SimplexTasks-12 包含 6 个受控合成场景与 6 个冻结预测器真实任务。

正文

View PDF HTML (experimental)

Abstract:Conformal prediction guarantees marginal coverage, but a single calibration threshold can still spread that coverage unevenly, over-covering easy regions and under-covering hard ones. SimplexUQ is, to our knowledge, the first benchmark and reproducible protocol for measuring this allocation problem on simplex-valued predictions; it compares existing conformal wrappers rather than proposing a new one. Its task suite, SimplexTasks-12, combines six controlled synthetic regimes with six frozen-predictor real tasks spanning class probabilities, topic mixtures, spectral abundances, cell-type fractions, age distributions, and emotion mixtures. Each comparison fixes the predictor, score, and response-free stratification map, varies only the wrapper, and reports marginal coverage, worst-stratum coverage, max disparity, and within-task radius and compute. Global calibration can look valid while failing badly: on CIFAR-10 it attains 0.900 marginal coverage but only 0.542 in the worst entropy stratum, and Mondrian calibration raises that stratum to 0.886 while reducing max disparity from 0.358 to 0.022. No wrapper dominates, however. Under smooth synthetic heterogeneity, several repairs are competitive; fixed-map analyses show that rankings depend on the evaluation groups and protocol; and in a 12-task comparison, Mondrian has lower disparity on its single target partition for all 12 tasks, whereas BatchMVP has lower disparity over overlapping groups on five. These are empirical comparisons, not new coverage guarantees. A controlled predictor-bias sweep shows that removing predictor bias only partly reduces global-threshold disparity. We release task cards, result provenance, permitted derived arrays, and rebuild instructions, and treat wrapper selection as a diagnostic comparison rather than a universal ranking.
Comments: Preprint. Code and data are available at this https URL and this https URL
Subjects: Machine Learning (cs.LG); Applications (stat.AP)
Cite as: arXiv:2610.00523 [cs.LG]
  (or arXiv:2610.00523v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00523

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Liang You [view email]
[v1] Wed, 30 Sep 2026 18:12:26 UTC (1,109 KB)

来源:arXiv:cs.LG · arxiv.org