跳到正文

#Anthropic

今日 5 条
今天10月1日周四
  1. IT之家(RSS)61

    谷歌推出 Gemini 4 Argon 旗舰模型,内部员工质疑其实战编码表现

    据彭博社报道,谷歌开始逐步推出旗舰模型 Gemini 4 Argon,先向一小批网络安全合作伙伴开放,之后优先面向付费订阅用户。谷歌称该模型多项基准测试靠前,安全测试成绩超过 OpenAI 的 Astra,但知情人士称其实际处理部分代码任务表现不佳,尤其前端设计能力参差不齐,且模型体量庞大、运行成本高。

9月30日周三
  1. Arena.ai62

    Arena 宣布 Claude Sonnet 5.5 (High) 可在 Direct Mode 中测试,限时 48 小时,至 10 月 2 日早 8 点(PT)结束,之后将继续在 Battle 和 Agent Mode 提供。引用的 @claudeai 内容称该模型是 Claude 5.5 家族第二款,比 Sonnet 5 快 30% 以上,多数工作成本降低最多 30%。

    引用Claude@claudeai

    Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

  2. Artificial Analysis 完整文章(网页)75

    Artificial Analysis 评测 GPT-6 Astra:与 Claude Fable 5.1 并列两大指数第一且成本更低

    Artificial Analysis 发布 GPT-6 Astra 基准测试报告,该模型在 Intelligence Index 得 53 分、Coding Agent Index 得 62 分,均与 Claude Fable 5.1 并列第一,且成本分别约为其 40% 和 60%。

    推荐理由:原文给出 GPT-6 Astra 在两大指数中的得分、成本和 token 效率数据,可据此比较它与 Claude Fable 5.1 的实际表现。

  3. Anthropic:The Institute(旗舰研究长文 · 网页)80

    Anthropic 发布 Claude Haiku 4.5,SWE-bench Verified 得分 73.3%

    Anthropic 发布 Claude Haiku 4.5,称其是最快、最具性价比的模型,在编码、计算机使用和 Agent 任务上媲美 Sonnet 4,SWE-bench Verified 得分 73.3%。

    推荐理由:官方给出性能、价格与可用渠道的关键数字,读者可据此评估它在低延迟和多智能体场景中的适用性。

  4. Anthropic:The Institute(旗舰研究长文 · 网页)82

    Anthropic 发布 Claude Sonnet 5.5,速度提升 30% 并降价最多 30%

    Anthropic 发布 Claude Sonnet 5.5,比 Sonnet 5 运行快 30%,典型负载成本最多低 30%,定价为每百万输入 token $2、输出 token $10,支持 1M 上下文窗口,可通过 Claude API、AWS、Google Cloud 和 Microsoft Foundry 使用。

    推荐理由:官方页面给出价格、token 用量与多客户评测数据,读者可以据此评估它在成本与性能间的取舍。

  5. Anthropic:The Institute(旗舰研究长文 · 网页)86

    Anthropic 发布 Claude Opus 5.5,运行成本比 Opus 5 低约 40%

    Anthropic 发布 Claude Opus 5.5,面向高强度编码和 AI 智能体的混合推理模型,具备 1M 上下文窗口,典型工作负载运行成本比 Opus 5 低约 40%。

    推荐理由:官方给出了定价、token 成本与多个客户实测数字,可帮助开发者评估升级到新 Opus 的成本与效率收益。

  6. Anthropic:The Institute(旗舰研究长文 · 网页)81

    Anthropic 发布 Claude Fable 5.1:面向编码与长时间智能体任务的新旗舰模型

    Anthropic 发布 Claude Fable 5.1,定位为其最强编码与知识工作模型,面向 Pro、Max、Team、Enterprise 用户及 Claude Platform、AWS、Google Cloud、Microsoft Foundry。

    推荐理由:官方页面给出定价、缓存降价、回退机制和客户实测引用,读者可以据此评估长时智能体与编码工作流的成本收益。

  7. Anthropic:The Institute(旗舰研究长文 · 网页)72

    Anthropic 发布 Claude Mythos 5.1,聚焦网络安全与生物研究并限制访问

    Anthropic 发布 Claude Mythos 5.1,是其面向网络安全和生物学研究的最强模型,目前仅向少量经审核的组织开放,定价为每百万输入 token $10、每百万输出 token $50。同底座的 Claude Fable 5.1 提供带护栏版本,生物护栏对良性请求的干预比 Fable 5 减少 85%,使用 Mythos 5.1 默认需接受 30 天数据保留政策。

    推荐理由:原文给出 Mythos 5.1 的受限开放方式、定价与 Fable 5.1 的护栏细节,读者可以据此了解高风险能力如何被分层开放。

  8. Artificial Analysis 完整文章(网页)78

    Claude Sonnet 5.5 登上 Artificial Analysis 智能指数第 2 名,token 用量创测量新高

    Anthropic 发布 Claude Sonnet 5.5,在 Artificial Analysis 智能指数得分 56(max 档较 Sonnet 5 提升 18 分),仅次于 Opus 5.5 (max)。

    推荐理由:原文给出 Sonnet 5.5 在智能指数、token 用量和成本上的具体数据,读者可据此对比其与 Opus 5.5 和 GPT-6 系列的实际取舍。

  9. Anthropic:The Institute(旗舰研究长文 · 网页)60

    Anthropic 发布 Claude Opus 5.5,运行成本比 Opus 5 低 40%

    Anthropic 于 9 月 22 日发布 Claude Opus 5.5,称其在多数工作上达到 Claude Fable 5.1 的水平,运行成本比 Opus 5 低 40%。同页还预告 9 月 10 日的威胁情报报告,介绍八个月内在多个行动中识别并阻断威胁行为者滥用 Claude 的案例,以及 2025 年以来恶意使用方式的演变。

  10. The Decoder:AI News(RSS)81

    OpenAI 发布 GPT-6.1 Sol,以五分之一成本接近 GPT-6.1 Astra

    OpenAI 发布 GPT-6.1 Sol,因安全问题推迟的旗舰 GPT-6.1 Astra 不按计划上线,Sol 作为更便宜的替代方案推出,OpenAI 称其在 agentic coding、计算机使用和办公任务上接近 Astra,成本约五分之一。

    推荐理由:原文汇总了 GPT-6.1 Sol 的定价与多项基准对比数据,读者可以据此判断它与 Astra、Claude 的成本性能权衡。

  11. Arena.ai78

    Arena 公布 Claude Sonnet 5.5 (High) 的实测结果,以 1699 分位列 Code Arena: WebDev 第 4,比 Sonnet 5 (High) 的 1540 分提升 159 分。

    引用Claude@claudeai

    Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

    推荐理由:Arena 的实测榜单数据显示该模型以约 1/5 的成本进入 WebDev 前四,读者可据此权衡性价比选型。

9月29日周二
  1. Thariq64

    Anthropic 发布 Claude Sonnet 5.5,是 Claude 5.5 家族的第二款模型,相比 Sonnet 5 明显升级,运行速度提升超过 30%,多数工作成本最高降低 30%。作者 Thariq 表示 Sonnet 与 Opus 5.5 让高阶智能更易获得,建议在构建工作流时尝试 Sonnet 5.5,以缓解 projects、claude tag 和 dynamic workflows 等抽象的 token 成本顾虑。

    引用Claude@claudeai

    Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

  2. Charlie Holtz71

    Charlie Holtz 转发 Anthropic 官方推文,宣布 Claude Sonnet 5.5 发布,是 Claude 5.5 家族的第二款模型。官方称其相比 Sonnet 5 是明显升级,运行速度提升超过 30%,多数工作场景成本最多降低 30%。

    引用Claude@claudeai

    Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

  3. ClaudeDevs79

    Anthropic 发布 Claude Sonnet 5.5,是 Claude 5.5 家族第二款模型,相比 Sonnet 5 更智能、更快 30% 以上,多数工作成本最多降低 30%。作者建议用于修复 bug、快速迭代功能等范围明确的日常任务,Claude Code 用量也因此更耐用。

    引用Claude@claudeai

    Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

    推荐理由:原文给出 Sonnet 5.5 相对 Sonnet 5 的速度、成本和适用场景,读者可据此判断是否切换日常 Claude Code 任务。

  4. Boris Cherny66

    Anthropic 发布 Claude Sonnet 5.5,是 Claude 5.5 家族第二款模型。相比 Sonnet 5,它速度提升超 30%,多数工作成本最高降低 30%。作者 Boris Cherny 附上用 Sonnet 5.5 在 Claude Code 中修复 bug 的视频演示。

    引用Claude@claudeai

    Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

    推荐理由:作者以 Claude Code 修复 bug 的视频演示新模型编码能力,可借此直观感受 Sonnet 5.5 相对上代的实际表现。

  5. Anthropic73

    Anthropic 官宣 Claude Sonnet 5.5 上线,是 Claude 5.5 家族的第二款模型。相比 Sonnet 5,速度提升超过 30%,多数工作的成本最高降低 30%。

    引用Claude@claudeai

    Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

    推荐理由:Anthropic 官方发布 Claude Sonnet 5.5,原文给出速度提升与降价幅度,读者可据此比较是否升级迁移。

9月28日周一
9月26日周六
  1. ClaudeDevs65

    Claude Devs 分析了 Opus 5.5 相比 Opus 5 的价格变化:输入和输出 token 便宜 20%,cache reads 便宜 60%。文章测算了 Claude Code 中一个任务的实际成本,并上线了一个计算器,用户可通过 /usage 自行运行,地址为 https://claude.dev/blog/what-a-task-costs-on-opus-5-5/。

    推荐理由:作者给出 Opus 5.5 相对 Opus 5 的价格降幅,并提供成本计算器,读者可自行估算 Claude Code 任务成本变化。

  2. Anthropic70

    Anthropic 科学博客发布文章,称 Claude 在单一提示词描述九圈问题后,在 Claude Science 中大体无监督运行数天,用 Dixon 等人开发的方法完成求解,总成本几千美元。该模型此前纪录为 SLAC 的 Lance Dixon 及合作者创下的八圈,Dixon 独立验证了这一结果,von Hippel 为博客撰写了这次经历。

    推荐理由:原文记录了 Claude 在九圈散射振幅计算中的求解过程与验证方式,读者可以了解学术级科学任务的可行路径。