跳到正文
arXiv:cs.LG· Idan Horowitz, Avigdor Gal·· 4 小时前AI 评分33

邻域平滑用于神经网络校准:基于图的训练时正则化方法

Neighborhood Smoothing for Calibration

AI 导读

研究者提出图平滑作为训练时校准的通用原则,通过惩罚表征空间中相邻样本预测分布间的 Jensen-Shannon 散度来改善模型校准,并推导出连接邻域预测散度与局部置信度变化及逐点校准误差传播的边界。该方法在标准校准基准上提升了预测质量,且与事后校准互补,在温度缩放后于八种图像和表格设置中的七种取得了所有评估训练时方法中最低的 NLL。

正文

View PDF HTML (experimental)

Abstract:Modern neural networks are often miscalibrated, with a tendency to overconfidence. Existing train-time calibration methods largely modify task losses or calibration penalties, leaving neighborhood structure in learned representations underexploited. We introduce graph smoothing as a general principle for train-time calibration, which encourages similar predictive distributions across neighboring samples in representation space. We analyze the effects of graph smoothing, deriving bounds that connect predictive divergence between neighboring samples to local confidence variation and to the propagation of pointwise calibration error, and characterize the conditions under which smoothing can or cannot improve calibration. In light of this analysis, we propose \modelNoSpace, a graph-based train-time regularizer that penalizes the Jensen--Shannon divergence between predictive distributions of neighboring samples. We present a thorough empirical analysis, showing that across standard calibration benchmarks, \model improves predictive quality, and the improvement is complementary to post-hoc calibration: after temperature scaling, \model attains the lowest NLL of all evaluated train-time methods in seven of the eight image and tabular settings. These findings demonstrate the value of graph smoothing over learned representations for neural network calibration.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.09020 [cs.LG]
  (or arXiv:2610.09020v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09020

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Idan Horowitz [view email]
[v1] Tue, 6 Oct 2026 19:16:15 UTC (1,498 KB)

来源:arXiv:cs.LG · arxiv.org