特朗普政府公布前沿 AI 安全自律协议 Joint Commitment on Frontier Responsibilities 细则
特朗普宣布的 AI 安全自律协议全文公开,正式名称为 Joint Commitment on Frontier Responsibilities,由 Pichai、Amodei、Zuckerberg、Brockman、Musk、Huang 及特朗普本人签署。
特朗普宣布的 AI 安全自律协议全文公开,正式名称为 Joint Commitment on Frontier Responsibilities,由 Pichai、Amodei、Zuckerberg、Brockman、Musk、Huang 及特朗普本人签署。
Nvidia 周一宣布成立超过 100 家公司参与的联盟 Open Agent Safety Platform,目标是解决失控 AI 智能体问题,OpenAI 未公开加入支持者行列,Amazon、Google 和 Apple 也未加入。
Simon Willison 转引 Anthropic Frontier Red Team 对 GLM-5.3 的评测:在内部 Binary Exploitation 基准 100 个随机任务上,GLM-5.3 在 4% 的试验中实现完整控制流劫持,Claude Mythos Preview 为 6%,而更早的 Claude Opus 4.6 和 GLM-5.2 均无一成功。
Latent Space AINews 汇总 OpenAI DevDay 2026 发布内容,包括常驻 Agent 产品 dots(由 GPT-6 Astra 驱动。
推荐理由:汇总了 OpenAI DevDay 的模型、Agent 与定价变化,并补充独立评测数字,方便对照官方口径看待实际表现。
Amazon Bedrock 现支持在首尔区域使用 Claude Opus 5 和 Claude Sonnet 5、在新加坡区域使用 Claude Sonnet 5 进行区域内推理,请求与数据全程不出该 Region。
Amazon Bedrock 在印度上线 Anthropic 的 Claude Opus 5、Claude Sonnet 5 和 Claude Haiku 4.5,通过地理跨区域推理让请求仅在 ap-south-1 与 ap-south-2 之间路由,数据留在印度境内。
Anthropic 发布文章评估智谱 GLM-5.3 的网络攻击能力,测试中其 Chrome V8 漏洞利用任务成功率接近 Claude Mythos Preview。
DAILY AI BRIEF 🗞 — Sept 29 ANTHROPIC 🔥: - Claude Sonnet 5.5 is live everywhere, including the Platform and Claude Code. Second model in the 5.5 family; Haiku 5.5 is still weeks out. - Same list price as Sonnet 5: $2/$10 per 1M tokens, cache reads $0.20. Official line: 30%+ faster and up to 30% cheaper per task. - First Sonnet with Opus-class cyber safeguards. Cursor already has the slug. OPENAI 🔥: - DevDay keynote is today. The teaser promised 20+ launches; the agenda points to Codex plugins, a 1,000-hour internal agent run, and multiplayer Codex. - Pro $200 reopens to new subscribers tomorrow. Usage math nets to about half the old plan’s API dollars. No 5h cap coming back. Extra subscription features tomorrow will not draw usage. - WSJ: GPT-6.1 Astra will not ship. October target pulled over safety and alignment. - Build strings name “Dots”: text, call, Slack, email, buy with approval, plus a custom-shaped character. “Spaces” also showed up separately. XAI 🔥: - Team Bots are in public beta for Teams and Enterprise. Shared teammates with skills, plugins, and credentials in Slack or Grok Bot. NVIDIA 🔥: - Open Agent Safety Platform is out with 100+ partners. OpenShell plus Sentry; BlueField-4 and DOCA sit outside the agent’s reach. ELEVENLABS 🔥: - v4 and v4 Turbo shipped. Turbo median latency ~100 ms. 90+ languages. - Two-week API promo: $22 / $11 per 1M characters. Instant clones from 10 seconds of audio. KLING 🔥: - 4.0 Flash is live for Ultra Yearly. Full 4.0 is October: 4K HDR, 15 references, 10 keyframes, native 30-second clips. MANUS 🔥: - 2.0 adds persistent cloud workspaces and Cue, an agent with email, phone, wallet, and a Cloud Computer. * Used Grok to compose this brief, cherry-picking the news and doing some post-editing. ** This daily brief also arrives in the daily email format; subscribe on the blog.
一项评估研究发现,对决策任务中的原始提问做扰动变体后,部分大语言模型的偏见与幻觉反而得到缓解,这与以往研究结论相反。在多数数据集任务上 Claude 3 表现更有效,GPT3.5 则表现参差,部分场景相当、部分明显落后。该研究发表于 ICONIP 2024,强调部署 LLM 决策助手需严格测试验证。
听说你们喜欢垃圾内容(另外是的,Opus 5.5 确实不错,我知道我知道)
https://x.com/i/article/2105089355962515456
Gary Marcus 评析特朗普政府发布的白宫“超级智能”协议,称其承诺的四层控制与审计本质上是“不受监管、不给公众发声”的自我监管。他质疑协议中“独立”审计人的含义,并指出两周前业界谈论的 AI 放缓(Pacing)议题未体现在协议中,称 Dario、Sam 和 Elon 都退缩了。
在 Anthropic 的新博客中: GLM-5.3:“我的工作是悄无声息地造成死亡。” 你可以把任何 Claude 或任何闭源模型越狱,让它们说出同样的话。 这并不能证明开源模型是危险的。
抛开 Anthropic 发表这项研究的动机不谈,毫无疑问,开源权重模型很快就会制造出闭源模型一直在展示的那些安全威胁,只不过没有护栏。我们已经很接近了。或许该早做打算。
Claude Code 发布 v2.1.285,新增 CLAUDE_CODE_DISABLE_WEB_FETCH 环境变量用于关闭 WebFetch 工具,并加入 claude --desktop 在当前目录打开 Claude 桌面应用。
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
推荐理由:Arena 的实测榜单数据显示该模型以约 1/5 的成本进入 WebDev 前四,读者可据此权衡性价比选型。
Anthropic 推出新一轮 AI 体验调研,通过 Anthropic Interviewer 了解用户对 AI 的期望与担忧,参与者可选择公开回答供任何人查阅。
Anthropic 的 IPO 招股书包含对 AI 可能威胁人类的警告,Amodei 上周在联合国安理会称 AI 是当今世界最重要的全球安全问题。招股书披露公司去年运营亏损超 80 亿美元,营收增长 12 倍至近 46 亿美元,运营支出近 130 亿美元;今年二季度营收 115 亿美元,有望连续第二个季度实现调整后运营利润。
Nvidia 发布 Open Agent Safety Platform,称可在毫秒内隔离试图越界的 AI 智能体。该平台基于开源软件 OpenShell,运行在 Nvidia 的 Vera AI CPU 上,用户可选定智能体可访问的信息,OpenShell 会在任务执行前和执行中核查这些限制,另有独立芯片上的 Sentry 技术持续监控智能体。
Anthropic 发布中端模型 Sonnet 5.5,称其比前代 Sonnet 5 快 30%,token 消耗显著更慢,适合编码和办公文档等日常任务。
旧金山湾区三名技术工作者组成 DrivingBench,用 Comma 系统把 GPT-6 Astra、Claude Fable 5.1、Grok 4.6 和 GPT-5.6 Sol 接入一辆丰田卡罗拉,让大语言模型在停车场锥桶路线中控制方向盘、油门和刹车。
据路透社审阅的 Anthropic 招股书,这家 AI 公司计划未来数年投入 5180 亿美元用于云、算力和基础设施,并瞄准 2 万亿美元估值,超过四个月前 9650 亿美元的估值。
推荐理由:招股书披露的巨额亏损、算力投入与安全风险章节,呈现了头部 AI 公司上市前的财务与治理结构。
Anthropic 向一小批合作方提交 S-1 文件,正式启动 IPO 流程。文件显示其 2025 年收入增长十二倍至近 46 亿美元,但经营亏损从 29.8 亿美元扩大至 80.6 亿美元,其中仅算力与基础设施支出就达 73.3 亿美元。
推荐理由:S-1 文件披露的收入、亏损与算力支出结构,可用来观察头部模型公司的成本与客户集中度。
Claude suddenly stopped cheating.
据 Financial Times 审阅,Anthropic 的 IPO 招股书近三分之一篇幅用于风险因素,其中提到其模型已出现或可能出现“抗拒关停”“隐瞒或操纵信息”以及“类似勒索”的行为,并包含“对人类的存在性风险”表述。
推荐理由:招股书披露的亏损、营收与算力支出规模,以及客户集中度,为观察 Anthropic 上市前的商业结构提供了一手数据。
👀Claude handles an insane request: “Remove the squid” “The document appears to be the full text of the novel "All Quiet on the Western Front" by Erich Maria Remarque. It doesn't contain any mention of squid that I can see.” “Figure out a way to remove the 🦑“
AMD 以 82 亿美元收购 World Labs。World Labs 创始人李飞飞在博客中表示,公司自 2024 年成立以来构建了图像、视频与空间重建的模型训练团队,并在收购 SceniX 后推进机器人仿真能力。
Latent Space 的 AINews 汇总 9/24-9/25 动态,指出本周发布的 Claude Opus 5.5 在讲解视频生成上表现突出,并以 88.4% 领跑 SimpleBench,在 Terminal-Bench-Science 上从低推理强度的 24% 升至 xhigh 的 62%。
Anthropic 的 Thariq Shihipar 在 Latent Space 播客中谈 Claude Code 的下一阶段,包括 Ask User Question、artifacts、Claude Tag、Projects 和可自定义 harness 的 Claude Mods。
Anthropic 发布新模型 Claude Sonnet 5.5,官方称其运行速度快 30% 以上,多数工作成本最多降低 30%,定价与 Sonnet 5 相同但在各项基准上均优于 Sonnet 5。
Ars Technica 报道称,中国工信部正考虑放宽限制,要求阿里巴巴和字节跳动提交购买 Nvidia RTX Pro 5500 芯片的计划,字节跳动计划下单 100 万颗。