跳到正文
原文
Anthropic:Transformer Circuits(可解释性研究)·· 16 小时前AI 评分35

Anthropic 可解释性研究:特征与上下文学习的新发现

Circuits Updates — September 2025 A small update on features and in-context learning.

AI 导读

Anthropic 可解释性团队在月度更新中指出,模型对跨语言相似文本的活跃特征重叠度(IoU)随样本长度增加而上升。通过对比英法段落首尾句,末句 IoU 显著高于首句,且无关语句的首句重叠度高于末句,支持"模型在更长上下文中构建更丰富表征"的解释,而非虚假激活所致。

来源:Anthropic:Transformer Circuits(可解释性研究) · transformer-circuits.pub