OpenAI:Alignment 研究博客(RSS)· Tom Dupre la Tour and Dan Mossing, in collaboration with the Interpretability team·· 2025-12-02AI 评分42
OpenAI 用稀疏自编码器潜在归因调试模型失准补全
Debugging misaligned completions with sparse-autoencoder latent attribution
AI 导读
OpenAI 提出用稀疏自编码器(SAE)的潜在归因(latent attribution)方法,替代此前的两步模型对比法,来定位与失准行为因果相关的 SAE 潜在特征。在涌现失准案例中,该方法选出的 top-100 Δ-attribution 潜在特征比 Δ-activation 特征更能通过激活引导让模型远离或走向失准,平均失准变化也更大;研究还用 GPT-5 对引导后的补全打分。
来源:OpenAI:Alignment 研究博客(RSS) · alignment.openai.com