arXiv:cs.LG· Jinhao Duan, Zicheng Liu, Zijie Liu, Kaidi Xu, Tianlong Chen·· 2 天前AI 评分39
LLM 如何偏离人类语言不确定性量化:ICML 2026 研究揭示置信语义差异
"very likely" Means "uncertain"? How LLMs Diverge from Humans in Linguistic Uncertainty Quantification
AI 导读
研究对比 LLM 与人类在语言不确定性量化上的差异,发现 LLM 为 "possible"、"likely" 等不确定性标记赋予的数值水平与人类存在显著偏差。作者提出一种基于优化的算法,直接从 LLM 输出中学习各标记的最优不确定性映射,无需重复采样即可实现标记级别的置信语义对比,揭示两者系统性置信差异。该工作已被 ICML 2026 接收。
正文
Abstract:Humans express uncertainty verbally via markers (e.g., "possible," "likely"), yet most LLM uncertainty quantification (UQ) relies on costing likelihood- or consistency-based signals. From a cognitive perspective, accurate verbal uncertainty reflects metacognitive monitoring, representing knowledge boundaries ("knowing that you don't know") to support regulation and information seeking. In this paper, we investigate how LLMs diverge from humans in verbal uncertainty quantification and whether verbal markers can reliably quantify LLM uncertainty. We curate a corpus of human uncertainty markers from psychology and decision-science literature and benchmark LLMs against it. We observe that LLMs encode verbal uncertainty with numerical levels that differ substantially from those of humans. We then introduce METHODNAME, a novel optimization-based algorithm that learns an optimal uncertainty profile over uncertainty markers directly from LLM outputs. By fitting a marker-uncertainty mapping to best explain empirical correctness, METHODNAME discovers how much probability mass each verbal marker should convey, rather than estimating uncertainty via repeated sampling. METHODNAME enables a direct, marker-level comparison of confidence semantics between humans and LLMs, disentangling mismatch and revealing systematic confidence disparities in verbal expressions.
| Comments: | ICML 2026 |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00083 [cs.LG] |
| (or arXiv:2610.00083v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00083 arXiv-issued DOI via DataCite |
Submission history
From: Zijie Liu [view email]
[v1]
Sun, 6 Sep 2026 19:31:45 UTC (897 KB)
来源:arXiv:cs.LG · arxiv.org