Meta 首席 AI 官 Alexandr Wang 表示,对齐是 AI 领域最开放的科学问题之一,目前没人知道确切解法,但已有一些思路。他提出的方案是"可扩展监督":用不同的 AI 去观察并约束更聪明的 AI,监督者 AI 必须与被监督模型同步进化,因此实验室需要构建"越来越聪明的监管智能体"。Meta 的 Muse 已采用这一思路,用独立的哨兵智能体检查主智能体的行为。
Alexandr Wang ( @alexandr_wang, Chief AI Officer at Meta): nobody knows how to solve alignment yet
"This is, I think, one of the most open questions scientifically in AI. I think nobody knows exactly the way to solve this problem, but there’s a few ideas."
His solution is "scalable oversight":
“as the AIs get smarter, we use a different set of AIs to observe what they’re doing and keep them in check.”
The watcher AIs have to improve along with the models they watch, so labs would need to build “smarter and smarter policing agents” too.
Meta’s Muse already uses a version of this, with a separate sentinel agent checking what the main agent does.
----
Full video on "Cleo Abram" YouTube channel, (link in comment)
来源:Rohan Paul · x.com