跳到正文
原文
Simon Willison 博客·· 1 天前AI 评分57

Anthropic Frontier Red Team 评测称 GLM-5.3 已具备二进制漏洞利用能力

Quoting Anthropic Frontier Red Team

AI 导读

Simon Willison 转引 Anthropic Frontier Red Team 对 GLM-5.3 的评测:在内部 Binary Exploitation 基准 100 个随机任务上,GLM-5.3 在 4% 的试验中实现完整控制流劫持,Claude Mythos Preview 为 6%,而更早的 Claude Opus 4.6 和 GLM-5.2 均无一成功。

来源:Simon Willison 博客 · simonwillison.net