跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Samira Alkaee Taleghan, Younghyun Koo, Andrew P. Barrett, Farnoush Banaei-Kashani·· 14 小时前AI 评分27

多专家区间标注下的不确定性感知学习

Uncertainty-Aware Learning from Multi-Expert Interval Targets

AI 导读

针对多专家区间标注中"标签内不精确"与"专家间分歧"两类不确定性,研究者提出一种保留各专家区间、并用 Beta 混合分布建模的方法,通过分解匹配将预测不确定性对齐到标签侧来源。在海冰浓度数据上,该模型较硬标签基线降低 MAE 31%,优于聚合、区间分布与区间回归基线。

正文

View PDF HTML (experimental)

Abstract:Many machine learning (ML) applications rely on expert labels, and qualified experts may provide different but plausible interpretations of the same observation. Such variation across expert labels may reflect genuine disagreement or ambiguity rather than annotation error. When individual experts additionally report intervals rather than exact values, the supervision contains two distinct sources of label uncertainty: within-label imprecision and between-expert variation. Existing methods treat these forms separately: multi-expert approaches collapse labels to a consensus, interval-target methods often yield a single prediction, and predictive-uncertainty methods rarely validate their uncertainty estimates against observed expert disagreement. To address this problem, we propose an approach that preserves individual expert intervals, separates within-label imprecision from between-expert variation, and validates the corresponding predictive uncertainty components. First, heterogeneous label vocabularies are harmonized into a common probabilistic label space, separating encoding differences from expert judgement. Second, individual label intervals are retained and modeled with a mixture of Beta distributions trained using a proper Cramér-distance objective, preserving distinct expert-reported labels. Third, we decompose predictive uncertainty into within-component, between-component, and model uncertainty, and evaluate whether these components correspond to within-label uncertainty, between-label uncertainty, and model error, respectively. Because this correspondence is not guaranteed, we introduce decomposition matching, which aligns the predictive components to their intended label-side sources. On sea-ice concentration the model reduces MAE by 31\% over hard labels and outperforms aggregation, interval-distribution, and interval-regression baselines.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.00102 [cs.LG]
  (or arXiv:2610.00102v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00102

arXiv-issued DOI via DataCite

Submission history

From: Samira Alkaee Taleghan [view email]
[v1] Tue, 8 Sep 2026 21:17:15 UTC (2,860 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org