跳到正文
原文
Anthropic:Transformer Circuits(可解释性研究)·· 13 小时前AI 评分59

Anthropic 开发三个对齐审计智能体并用审计游戏验证其能力与局限

Automated Auditing A note on using agents to perform automated alignment audits, including using interpretability tools.

AI 导读

Anthropic 开发了三个自主执行对齐审计任务的智能体,并在含植入缺陷的目标模型上验证。调查智能体在真实条件下以 13% 的成功率解开 Marks et al. 审计游戏,通过外层智能体循环聚合多次调查提升到 42%;评估智能体构建的行为评估在 88% 的运行中区分出有无植入行为的模型;广度优先红队智能体发现 10 个植入行为中的 7 个。

来源:Anthropic:Transformer Circuits(可解释性研究) · alignment.anthropic.com