跳到正文
arXiv:cs.AI· Robert Graham, Yariv Barsheshat, Phil Blandfort, Sabri Alouache·· 3 小时前

模型一致性研究:窄域微调模型在重采样下自相矛盾

An Investigation of Model Coherence: Narrow Finetunes Contradict Themselves Under Resampling

AI 导读

研究者提出一种基于重采样矛盾的高特异性一致性指标,能识别现有输出方差方法忽略的歧义与无关两类干扰因素,并构建了175个问题数据集。测试发现窄域微调模型得分较差,文献中的模型有机体表现出身份混淆、内省失败和合理化等严重问题。这表明窄域微调引入的缺陷可能限制其用于研究一致性失准行为的价值。

正文

View PDF HTML (experimental)

Abstract:A large body of research measures model coherence based on output variance without adequately considering competing causes. We identify two such causes, ambiguity and indifference, and we introduce a set of 175 questions where contradicting answers cannot easily be explained by either. We then measure incoherence in terms of contradictions when resampling answers to the same question. In contrast to other methods our metric has high specificity, and only ranks models as incoherent when the issues are glaring. Even so, we find narrow finetunes score poorly. Inspecting inconsistencies flagged by our method, we find that model organisms from the literature display severe issues such as identity conflation, introspection failures and rationalizations. These findings suggest that the pathologies induced by narrow finetuning may limit what these models can tell us about coherent misaligned behaviour.
Comments: 18 pages, 5 figures, 4 tables. Code and dataset: this https URL
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.12129 [cs.AI]
  (or arXiv:2610.12129v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.12129

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yariv Barsheshat [view email]
[v1] Thu, 8 Oct 2026 15:22:00 UTC (321 KB)

来源:arXiv:cs.AI · arxiv.org