跳到正文
arXiv:cs.LG· Eug\`ene Berta, Sacha Braun, Francis Bach, Michael I. Jordan, David Holzm\"uller·· 3 小时前

V-ECE:估计通用期望校准误差

V-ECE: Estimating General Expected Calibration Errors

AI 导读

研究者提出变分估计器 V-ECE,无需分箱或聚类,可估计基于通用凸散度(含 L_p 距离)的校准误差,并给出二分类与多分类场景的闭式损失。该方法通过梯度提升拟合温度缩放的残差来估计超额风险,期望上给出真实校准误差的下界。在由真实分类器构建、已知真实 CE 的半合成基准上,V-ECE 在各类校准误差下均为最准确的二分类估计器之一,并显著优于所有多分类估计器。

正文

View PDF HTML (experimental)

Abstract:In probabilistic classification, calibration error (CE) measures the average divergence of predicted probabilities $f(X)$ from $\mathbb{P}(Y|f(X))$, the true class distribution for that predicted probability. While being a useful diagnostic tool, it is hard to estimate: popular binning-based estimators are often inconsistent and scale poorly beyond two classes. Recent work rewrites the CE as the excess risk of a model compared to the best recalibration of its own predictions, measured with a proper loss. However, this only works for Bregman-divergence-based calibration errors like the squared error, excluding the more popular $L_1$-distance-based CE. We show that using prediction-dependent proper scores can alleviate this restriction, allowing us to estimate CEs with general convex divergences, including $L_p$ distances with closed-form losses in the binary and multiclass settings. To estimate the excess risk, we introduce a more accurate recalibrator that fits a residual to temperature scaling with gradient boosting. The resulting variational estimator, V-ECE, needs no bins or clusters and lower-bounds the true calibration error in expectation. On a benchmark of semi-synthetic tasks built from real classifiers, with known true CE, V-ECE is among the most accurate binary estimators for every calibration error and significantly outperforms all multiclass estimators. Our results are accompanied by additional theory on $L_p$ CE, estimator bias, and over- or under-confidence estimation.
Comments: Re-worked version with a new metric benchmark and new mathematical results on the bias of the estimator
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG)
Cite as: arXiv:2602.24230 [stat.ML]
  (or arXiv:2602.24230v2 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2602.24230

arXiv-issued DOI via DataCite

Submission history

From: Eugène Berta [view email]
[v1] Fri, 27 Feb 2026 17:56:52 UTC (100 KB)
[v2] Thu, 8 Oct 2026 16:04:35 UTC (1,522 KB)

来源:arXiv:cs.LG · arxiv.org