Anthropic:Transformer Circuits(可解释性研究)·· 16 小时前AI 评分35
Anthropic 可解释性研究:特征与上下文学习的新发现
Circuits Updates — September 2025 A small update on features and in-context learning.
AI 导读
Anthropic 可解释性团队在月度更新中指出,模型对跨语言相似文本的活跃特征重叠度(IoU)随样本长度增加而上升。通过对比英法段落首尾句,末句 IoU 显著高于首句,且无关语句的首句重叠度高于末句,支持"模型在更长上下文中构建更丰富表征"的解释,而非虚假激活所致。
来源:Anthropic:Transformer Circuits(可解释性研究) · transformer-circuits.pub