OpenAI:Alignment 研究博客(RSS)· Marcus Williams, Cameron Raymond and Micah Carroll, in collaboration with the Safety Oversight team·· 2025-12-19精选AI 评分66
OpenAI 发布生产评测方法,用去标识 ChatGPT 流量规避评估感知并预测模型未对齐行为
Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluations
AI 导读
OpenAI Alignment 团队发布生产评测(production evaluations)方法,用去标识的 ChatGPT 生产流量剥离最终回复后重采样新模型输出,再用 LLM 监测器发现新的未对齐行为并估计其发生率。
推荐理由
OpenAI 团队详述了用去标识生产流量构建对齐评测的完整流程、验证数据和局限,读者可以据此理解生产评测如何降低评估感知。
来源:OpenAI:Alignment 研究博客(RSS) · alignment.openai.com