跳到正文
原文
METR:Research(网页)·· 13 小时前AI 评分57

METR 研究:即使 CoT 不忠实,仍可用于监控模型行为

CoT May Be Highly Informative Despite “Unfaithfulness” August 8, 2025 Recent work from Anthropic and others claims that LLMs' chains of thoughts can be “unfaithful”. These papers make an important point: you can't take everything in the CoT at face value. As a result, people often use these results to conclude the CoT is useless for analyzing and monitoring AIs. Here, instead of asking whether the CoT always contains all information relevant to a model's decision-making in all problems, we ask if it contains enough information to allow developers to monitor models in practice. Our experiments suggest that it might. Read more

AI 导读

METR 发布研究,复现并改进 Anthropic 的 CoT 忠实性评估,测试 Claude Sonnet 3.7、Claude Opus 4 和 Qwen 3 235B。

来源:METR:Research(网页) · metr.org