跳到正文
原文
Anthropic:Transformer Circuits(可解释性研究)·· 13 小时前AI 评分34

Anthropic 可解释性团队 Circuits Updates:2023 年 7 月研究进展

Circuits Updates — July 2023 A collection of small updates from the Anthropic Interpretability Team.

AI 导读

Anthropic 可解释性团队发布 2023 年 7 月 Circuits Updates,汇总多项尚在萌芽阶段的研究想法。团队重新审视了 Superposition、Memorization 与 Double Descent 中的中间数据区间,认为此前疑似反例的现象实为低维优化失败造成的假象,并发现模型在约 1k 至 500k 样本间会从记忆单点转向记忆相关数据簇。

来源:Anthropic:Transformer Circuits(可解释性研究) · transformer-circuits.pub