arXiv:cs.CL· Dianqiao Lei, Kevin Qinghong Lin, Pan Lu, Philip Torr, James Zou·· 4 小时前
LLM 告诉它别想却停不下来:研究揭示"思考惯性"
Thinking Inertia: LLMs Keep Thinking When Told Not To
AI 导读
研究发现 LLM 存在"思考惯性"(Thinking Inertia):即使明确关闭思考模式或提示不要推理,模型仍会持续输出显式推理,且答案空间越开放越明显。
正文
Abstract:Large Language Models (LLMs) increasingly ship with explicit "thinking modes", yet their counterpart, "no-thinking", has received far less attention. We study LLMs' no-thinking behavior along two axes. a. How to measure no-thinking? Prior work typically defines no-thinking through proxies such as a disabled thinking mode or the absence of long traces. These proxies are unreliable: disabled thinking modes may still emit reasoning, while long traces may contain filler rather than genuine inference. We instead normalize each response into a pre-answer trace and final answer, and evaluate it at three levels: (i) Empty-Thinking Rate for strict answer-only compliance; (ii) instruction-aware Question-Pre-answer Relevance for similarity between the question and pre-answer trace; and (iii) LLM-as-judge Explicit Inference Rate for visible explicit inference. Together, these metrics distinguish answer-only output, relevant but non-inferential text, and explicit inference. b. How does no-thinking vary across tasks and models? We evaluate six prompting interventions on six LLMs across Boolean, multiple-choice, and open-ended questions. We find that explicit no-think controls cannot reliably eliminate visible inference. Models instead exhibit "Thinking Inertia": explicit inference persists even under strict controls and becomes more prevalent as the answer space opens. Accuracy remains stable on Boolean and multiple-choice tasks, whereas open-ended tasks reveal a trade-off between answer-only compliance and task accuracy. Rewriting the same questions across answer spaces shows that supplying candidate answers makes answer-only responses easier to produce. These findings establish no-thinking as a non-trivial capability: stopping explicit reasoning cannot be assumed from model settings or instructions alone and deserves systematic evaluation alongside reasoning ability.
| Comments: | Accepted by NeurIPS 2026. Website: this https URL GitHub: this https URL |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.11765 [cs.CL] |
| (or arXiv:2610.11765v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11765 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Qinghong Lin [view email]
[v1]
Thu, 8 Oct 2026 11:49:37 UTC (1,151 KB)
来源:arXiv:cs.CL · arxiv.org