Anthropic:Transformer Circuits(可解释性研究)·· 12 小时前AI 评分39
Anthropic 可解释性团队 Circuits Updates(2024 年 2 月):字典学习中的死亡特征与改进
Circuits Updates — February 2024 A collection of small updates from the Anthropic Interpretability Team.
AI 导读
Anthropic 可解释性团队发布 2024 年 2 月 Circuits Updates,聚焦稀疏自编码器字典学习中的"超低密度特征"问题,认为这类特征因 L1 正则惩罚在找到有用方向前就被杀死。
来源:Anthropic:Transformer Circuits(可解释性研究) · transformer-circuits.pub