跳到正文
arXiv:cs.CL· David I. Atkinson, Dillon Plunkett, David Bau·· 6 小时前AI 评分48

如何从模型内部识别真实的内省式自我报告

Identifying Introspection From the Inside

AI 导读

研究者通过低秩适配器训练模型按隐含线性偏好函数为虚构角色做决策,发现在隐式决策任务上持续微调可让模型在无显式自我报告监督下,自发涌现出对所学偏好的准确自我报告。

正文

View PDF HTML (experimental)

Abstract:Large language models make claims about themselves that are both consequential and increasingly difficult to verify from behavior alone. How can we distinguish plausible confabulations from genuine introspection? In this paper, we identify mechanistic signatures of faithful self-report in a controlled setting. Using low-rank adapters, we train models to make decisions on behalf of fictitious characters, according to latent linear preference functions. We find sustained fine-tuning on an implicit decision task can lead to the emergence of accurate self-reporting of models' learned preferences, even without explicit self-report supervision. We ask two research questions about this emergent phenomenon. First: is the emergence of accurate self-reporting accompanied by a measurable structural change in the model? Weight ablations and frozen-layer experiments together indicate that preference representations shift to earlier layers over training, consistent with the hypothesis that faithful self-report requires preferences to be located where pre-existing verbalization mechanisms can access them. Second: can these structural differences distinguish faithful models from unfaithful ones? Using attribution patching, we find that faithful models exhibit significantly higher attribution similarity between the decision-making and self-report tasks -- a mechanistic signature of faithful self-report that does not require us to understand the content of the report itself. Previous work on self-report has observed behaviorally that models can be faithful or unfaithful; our work proposes that, at least in our restricted setting, it is possible to distinguish between the two patterns of computation by examining the structure of the networks themselves.
Comments: Published at COLM 2026. Project page: this https URL
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.07186 [cs.CL]
  (or arXiv:2610.07186v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.07186

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: David Atkinson [view email]
[v1] Mon, 5 Oct 2026 18:06:48 UTC (335 KB)

来源:arXiv:cs.CL · arxiv.org