arXiv:cs.LG· Eug\`ene Berta, Sacha Braun, Francis Bach, Michael I. Jordan, David Holzm\"uller·· 3 小时前
V-ECE:估计通用期望校准误差
V-ECE: Estimating General Expected Calibration Errors
AI 导读
研究者提出变分估计器 V-ECE,无需分箱或聚类,可估计基于通用凸散度(含 L_p 距离)的校准误差,并给出二分类与多分类场景的闭式损失。该方法通过梯度提升拟合温度缩放的残差来估计超额风险,期望上给出真实校准误差的下界。在由真实分类器构建、已知真实 CE 的半合成基准上,V-ECE 在各类校准误差下均为最准确的二分类估计器之一,并显著优于所有多分类估计器。
正文
Abstract:In probabilistic classification, calibration error (CE) measures the average divergence of predicted probabilities $f(X)$ from $\mathbb{P}(Y|f(X))$, the true class distribution for that predicted probability. While being a useful diagnostic tool, it is hard to estimate: popular binning-based estimators are often inconsistent and scale poorly beyond two classes. Recent work rewrites the CE as the excess risk of a model compared to the best recalibration of its own predictions, measured with a proper loss. However, this only works for Bregman-divergence-based calibration errors like the squared error, excluding the more popular $L_1$-distance-based CE. We show that using prediction-dependent proper scores can alleviate this restriction, allowing us to estimate CEs with general convex divergences, including $L_p$ distances with closed-form losses in the binary and multiclass settings. To estimate the excess risk, we introduce a more accurate recalibrator that fits a residual to temperature scaling with gradient boosting. The resulting variational estimator, V-ECE, needs no bins or clusters and lower-bounds the true calibration error in expectation. On a benchmark of semi-synthetic tasks built from real classifiers, with known true CE, V-ECE is among the most accurate binary estimators for every calibration error and significantly outperforms all multiclass estimators. Our results are accompanied by additional theory on $L_p$ CE, estimator bias, and over- or under-confidence estimation.
| Comments: | Re-worked version with a new metric benchmark and new mathematical results on the bias of the estimator |
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG) |
| Cite as: | arXiv:2602.24230 [stat.ML] |
| (or arXiv:2602.24230v2 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2602.24230 arXiv-issued DOI via DataCite |
Submission history
From: Eugène Berta [view email]
[v1]
Fri, 27 Feb 2026 17:56:52 UTC (100 KB)
[v2]
Thu, 8 Oct 2026 16:04:35 UTC (1,522 KB)
来源:arXiv:cs.LG · arxiv.org