跳到正文
arXiv:cs.CL· Tobias Braun, Nils Loose, Alexander Herzog, Virginia Ceccatelli, Marcus Rohrbach, Thomas Eisenbarth, Lorenzo Cavallaro·· 3 小时前AI 评分46

U-Space:揭示语言模型中的不确定性何时以及为何产生

U-Space: Uncovering When and Why Uncertainty Arises in Language Models

AI 导读

研究者提出 U-Space,一个低维子空间,用于让语言模型演化中的不确定性变得可测量、可解释。其 U-Lens 将每个 token 状态投影到由“怀疑”与“确定”语义锚点构成的正交基上,生成 token 级不确定性图,并可聚合为标量置信分数,无需正确性标签、重复生成或训练。在推理基准上,该置信分数在标准与长度控制评测下均优于已有基线,且比监督式估计器迁移更可靠。

正文

View PDF HTML (experimental)

Abstract:Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer remains difficult because language models can present incorrect conclusions with fluent explanations and an authoritative tone. Uncertainty quantification seeks to address this disconnect by estimating the reliability of individual predictions. However, many existing methods require repeated generations or separately trained components, and their scalar estimates do not reveal where uncertainty arises or how it evolves during reasoning. Recent work has also shown that generation length can be strongly associated with uncertainty estimates and correctness, raising the question of how much of an estimator's predictive power comes from uncertainty-specific information rather than output length alone. Mechanistic interpretability offers a way to address these limitations by connecting human-interpretable concepts to intermediate model states. Building on this capability, we introduce the U-Space, a low-dimensional subspace that makes a model's evolving uncertainty measurable and interpretable. We identify semantic anchors for doubt and certainty, map their unembedding directions back into the residual space, and combine their contrasts into an orthogonal basis. The U-Lens projects each token state onto these basis vectors, yielding an interpretable token-level uncertainty map that can be inspected directly or aggregated into a scalar uncertainty score. Our approach requires no correctness labels, repeated generations, or training. Across reasoning benchmarks, its confidence score outperforms established baselines under both standard and length-controlled evaluation and transfers more reliably than supervised estimators. Code: this https URL.
Comments: Code: this https URL
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.09087 [cs.CL]
  (or arXiv:2610.09087v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.09087

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Tobias Braun [view email]
[v1] Tue, 6 Oct 2026 20:40:29 UTC (1,309 KB)

来源:arXiv:cs.CL · arxiv.org