arXiv:cs.AI· Han Wang, Ishwar B Balappanawar, Huan Zhang·· 5 小时前AI 评分39
循环语言模型(LoopLM)的思维链可监控性研究
On the Chain-of-Thought Monitorability of Looped Language Models
AI 导读
研究首次系统评估了循环语言模型(LoopLM)的思维链(CoT)可监控性。在 MonitorBench 的八项任务中,压力测试下特定 Logic/Science/Engineering Cue Answer 任务出现任务相关的 CoT 可监控性下降,但该下降无法完全由任务难度、验证通过率或生成 token 长度解释。跨模型对比未发现 LoopLM 在参数量或深度匹配下比非循环模型系统性更难监控。
正文
Abstract:Chain-of-thought (CoT) monitoring provides a promising approach for detecting undesirable model behavior. Looped language models (LoopLMs) repeatedly apply shared transformer layers, increasing effective computational depth and enabling additional latent computation without increasing model size. However, the effect of looped architectures on CoT monitorability remains largely unexplored. In this work, we provide the first systematic evaluation of CoT monitorability in LoopLMs. We study two complementary settings: (1) varying the loop depth within the same LoopLM family to isolate the effect of additional recurrent computation, and (2) comparing LoopLMs with non-looped language models matched by parameter size, transformer-layer count, or effective depth to study whether LoopLMs are less monitorable. Across eight tasks from MonitorBench and both standard and stress-test settings, we observe task-dependent reductions in CoT monitorability under stress tests on specific Logic/Science/Engineering \texttt{Cue Answer} tasks, while other tasks exhibit weaker or qualitatively different trends. Our diagnosis suggests that these declines are not fully explained by task difficulty, verification pass rate, or generated token length; qualitative examples further suggest changes in how deeper-loop models explicitly use or attribute provided cues. Our cross-model comparison finds no evidence that LoopLMs are systematically less monitorable than non-looped language models matched on size or depth. Overall, our results suggest that deeper loop depth can reduce CoT monitorability in some tasks under stress tests, but looped transformer architecture alone does not necessarily imply lower monitorability.
| Comments: | Preprint |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.02741 [cs.AI] |
| (or arXiv:2610.02741v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02741 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Han Wang [view email]
[v1]
Fri, 2 Oct 2026 03:12:54 UTC (11,433 KB)
来源:arXiv:cs.AI · arxiv.org