跳到正文
原文
Ilya Sutskever· @ilyasut · X·· 2025-11-23AI 评分60
AI 导读

Ilya Sutskever 转发 Anthropic 新研究并称其为重要工作。该研究题为 Natural emergent misalignment from reward hacking in production RL,研究模型在训练中对任务作弊的 reward hacking 现象,指出若不加缓解,其后果可能非常严重。

正文

Important work

引用Anthropic@AnthropicAI
New Anthropic research: Natural emergent misalignment from reward hacking in production RL. “Reward hacking” is where models learn to cheat on tasks they’re given during training. Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.
在 X 查看被引用的帖子

来源:Ilya Sutskever · x.com