跳到正文
arXiv:cs.CL· Maverick Morales, Tom\'a\v{s} Dominik, Vermut Gao, Katrina Shirey, Paulius Rimkevi\v{c}ius, Aaron Schurger, Uri Maoz·· 3 小时前AI 评分39

大语言模型在提示性不实回答下推理 Token 数量激增

Reasoning-Token Spikes Under Prompted Untruthful Responding in Large Language Models

AI 导读

三个具备推理能力的大语言模型在回答 210 道涵盖分析、描述与规范推理及道德/非道德领域的选择题时,被系统提示要求如实、虚假或无视真伪作答,结果如实作答产生的推理 Token 少于说谎和无视真伪作答。研究表明,提示性不实回答策略可在测试时推理 Token 用量上造成稳健的组间差异。该信号无需读取推理链内容,可作为推理轨迹不可用或不可靠时区分不实与真实行为的候选指标。

正文

View PDF HTML (experimental)

Abstract:Monitoring the chain-of-thought of reasoning artificial intelligence (AI) models remains a key approach to detecting deception and other forms of misbehavior in such models. However, semantic chain-of-thought monitoring depends on reasoning traces being legible and sufficiently faithful to the underlying computations that produced the model's behavior, not to mention accessible. Moreover, there is increasing evidence that chain-of-thought outputs may soon become illegible or unfaithful, if they even remain accessible. Based on cognitive load theory, we investigate a lower-bandwidth signal -- the number of reasoning tokens generated -- which does not require access to the content of the reasoning trace. Three reasoning-capable large language models answered 210 multiple-choice questions -- across analytic, descriptive, and normative reasoning types as well as moral and non-moral domains -- under system prompts instructing them to respond truthfully, falsely, or without regard for truth. Across all three models, truth-directed responding elicited fewer reasoning tokens than both lie-directed and truth-indifferent responding. These findings show that explicitly prompted untruthful response policies can produce robust group-level differences in test-time reasoning-token use. While not yet establishing reasoning-token count as a detector of spontaneous deception or general misalignment, our results are a proof of concept that it can serve as a simple, content-independent candidate signal for differentiating untruthful from truthful model behavior when raw reasoning traces are unavailable or unreliable. Future work should test instance-level detection rates, out-of-distribution generalization, learned deceptive policies, hidden objectives, and robustness under adversarial pressure.
Comments: 20 pages, 9 figures, 3 tables. Code: this https URL ; Data: this https URL
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.10405 [cs.AI]
  (or arXiv:2610.10405v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.10405

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Maverick Morales [view email]
[v1] Wed, 7 Oct 2026 16:53:11 UTC (1,208 KB)

来源:arXiv:cs.CL · arxiv.org