Anthropic:Transformer Circuits(可解释性研究)·· 13 小时前AI 评分38
Anthropic 可解释性团队 Circuits Updates:TopK 与 Gated SAE 对比研究
Circuits Updates — June 2024 A collection of small updates: topk and gated SAE investigation.
AI 导读
Anthropic 可解释性团队复现了 TopK SAE 与 Gated SAE 相对标准 L1 惩罚 SAE 的改进,确认两者在 L0/MSE 权衡上均有显著提升且表现相近。在 100 个特征的盲评中,TopK 与 Gated SAE 的特征可解释性未见明显退化,但 TopK 在高密度区间特征更多、平均激活变化更大。团队尚未实现 Gao 等人用于防止死特征的 Ghost Grads 变体。
来源:Anthropic:Transformer Circuits(可解释性研究) · transformer-circuits.pub