跳到正文
原文
xAI:News(网页)·· 12 小时前AI 评分43

LatchBio 独立评测:Grok 4.6 生物安全监测与对抗性生物任务表现超越所有前沿模型

Biosecurity at the frontier LatchBio evaluated Grok's performance on biosecurity monitoring and adversarial biological tasks. They found that Grok 4.6 detects and refuses dangerous queries more reliably than any other frontier system. Sep 1, 2026

AI 导读

LatchBio 对 Grok 4.6 的独立评测显示,其在 BioSecBench-Refusal 上拒绝伪装危险任务的表现优于所有受测前沿模型,是唯一在红队拒绝率(59.2%)与常规任务完成率(64.8%)两项指标上均超过 50% 的系统,跨不同 agent harness 平均得分 62.1%。

来源:xAI:News(网页) · x.ai