跳到正文
原文
Anthropic:Transformer Circuits(可解释性研究)·· 12 小时前AI 评分46

Anthropic 详解可解释性工具 Garçon:面向大模型的基础设施

Garcon A description of our tooling for doing interpretability on large models. note, infrastructure

AI 导读

Anthropic 公开其可解释性研究基础设施 Garçon,用于在超出单节点规模的大模型上做机制可解释性实验。用户可通过 `python -m garcon.launch` 启动服务器并连接模型,借助 probe points 与 probe 函数在任意层读取、修改或消融内部激活,并支持前向与反向传播。服务器按连接保存数据,空闲 1 小时后自动关闭。

来源:Anthropic:Transformer Circuits(可解释性研究) · transformer-circuits.pub