跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Sharath M Shankaranarayana, Davor Runje, Jan Jannink·· 9 小时前AI 评分38

黑盒决策模型 Jev 的自我知识审计:置信度校准良好但无法识别知识缺失

Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model

AI 导读

研究对决策模型 Jev 在 15 个以上公开数据集和 6 类生成任务上开展受控审计,发现其置信度在熟悉的封闭选择题上校准良好,却无法反映知识缺失:在无答案相关信息时对显著选项给出高达 0.80 的置信度,在超出知识边界的新闻上置信度超出准确率 0.21–0.33,且用早期月份数据重校准也无法弥合。

正文

View PDF HTML (experimental)

Abstract:Decision models return probabilities intended for routing, abstention and automated action. Calibration makes those probabilities useful on average, but does not establish whether low confidence reflects chance or missing knowledge, nor whether confidence falls when a model moves beyond what it knows. We audit this distinction in Jev, a decision model, with over 15 public datasets and 6 generated task families, with paired interventions that vary the information supplied for a fixed item. Jev's confidence is calibrated on familiar closed-choice tasks but fails as an indicator of missing knowledge: with no answer-relevant information it assigns up to 0.80 to a salient option, and on news beyond an observed knowledge boundary it exceeds accuracy by 0.21--0.33, a gap that recalibration on earlier months does not close. Targeted yes/no questions give sharper readouts of the case: whether an outcome is settled (AUROC 1.00) and whether the evidence suffices (0.95, against 0.85 for confidence on the same items). Asking whether Jev knows the answer appears to flag fabricated entities and post-boundary news (0.91), but with realistic names or with dates removed it shows no advantage over answer uncertainty. Black-box knowledge audits therefore need explicit controls for surface cues. Code: this https URL.
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.01006 [cs.AI]
  (or arXiv:2610.01006v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.01006

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sharath M. Shankaranarayana Mr [view email]
[v1] Thu, 1 Oct 2026 03:53:12 UTC (894 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org