Dan Hendrycks· @hendrycks · X·· 2026-08-22AI 评分43
AI 导读
更多发现表明,可解释性工具很脆弱,甚至不如简单的基线方法。
正文
More findings that interpretability tools are fragile or worse than simple baselines.
A good explanation of a model's behavior should help you make predictions in related situations. We turn this into an eval, with thousands of real behaviors found in the wild. Can interp tools help here? On average, no. 🧵在 X 查看被引用的帖子
来源:Dan Hendrycks · x.com