Anthropic:Transformer Circuits(可解释性研究)·· 13 小时前AI 评分39
Anthropic 可解释性研究:逆向工程 Transformer 的参数级练习
Exercises Some exercises we've developed to improve our understanding of how neural networks implement algorithms at the parameter level. note, exercises
AI 导读
Anthropic 的 Transformer Circuits 团队发布了一组练习,用于理解神经网络如何在参数层面实现算法,作为其逆向工程 Transformer 数学框架的补充材料。练习要求手动写出注意力头的 W_Q、W_K、W_V、W_out 权重矩阵,涵盖虚拟注意力头、归纳头(指针算术版与前一 token K-Composition 版)等实现,并附有解答。
来源:Anthropic:Transformer Circuits(可解释性研究) · transformer-circuits.pub