跳到正文
原文
Anthropic:Transformer Circuits(可解释性研究)·· 16 小时前AI 评分43

Anthropic 可解释性研究:用下游连接预测哪些特征会操控模型行为

Circuits Updates — May 2026 A short update on understanding features through downstream connections.

AI 导读

Anthropic 可解释性团队提出用 TWERA 加权的下游连接来区分外观相似的特征:两个在“草是什么颜色”提示上同样激活、top unembeds 都指向 green 的 CLT 特征,只有 Feature B 被抑制才会让答案从 Green 变成 Red。实验收集 10 组各 3–5 个候选特征,让 Opus 4.7 按操控可能性排序,加入 TWERA 下游特征信息后预测能力小幅提升。

来源:Anthropic:Transformer Circuits(可解释性研究) · transformer-circuits.pub