跳到正文

#Anthropic

今日 4 条
9月23日周三
  1. Claude Code:GitHub Releases(RSS)65

    Claude Code v2.1.280 发布:新增 Claude Opus 5.5 默认模型

    Claude Code 发布 v2.1.280,新增 Claude Opus 5.5(claude-opus-5-5)并设为默认 Opus 模型,支持 1M 上下文,价格 $4/$20 per Mtok、缓存读取 $0.20/Mtok;Pro 和 Team Standard 计划默认模型也从 Sonnet 改为 Opus。

    推荐理由:原文列出该版本新增 Claude Opus 5.5 默认模型、MCP 描述长度可配置等改动,读者可对照修复清单决定是否升级。

9月22日周二
9月21日周一
  1. Peter McCrory24

    大体同意。一些实际启示: (1) 优先做能用新数据定期更新的分析 (2) 公开地做研究(根据新证据修正观点) (3) 承认不确定性;做出可证伪的预测 (4) 认真且谦逊

    引用Alex Imas@alexolegimas

    A few (personal) thoughts on reading empirical AI papers on the economy. Economists have gotten used to reading papers with super clean identification, arguing about the validity of an instrument, making sure parallel trend assumptions are satisfied. This is what gets you into a top journal, and it is *very* important research (no question here). But it also takes years and sometimes decades to get these types of papers right---people often don't find a good instrument to answer a specific causal question decades after the natural experiment. We will eventually have this type of research for AI as well, and it is absolutely necessary. But right we also need signals *right now*, even if they are noisier than what we are used to. We need papers where we can trust that researchers did their best methodologically, while at the same time acknowledging that the space is moving way too fast to wait for perfect identification. This will allow us to accumulate enough signals, coming at the same question using different angles, for example, to say "yes, X is likely happening in the economy". The AI exposure and early career hiring papers are a good example of this. There is no silver bullet paper with super clean identification. But at this point we have several independent teams reaching the same general conclusion, enough where we can say "there seems to be a slow down in AI-exposed, early career hiring."

9月19日周六
  1. Gary Marcus:The Road to AI We Can Trust(RSS)33

    Gary Marcus 批评 Dario Amodei 七天内三度失信

    Gary Marcus 发文列举 Dario Amodei 在七天内损害自身公信力的三种做法:其一是让与 Anthropic 关系密切的 METR 和已有业务往来的 Accenture 充当独立监督方;其二是 Anthropic 正筹备自建湿实验室,却缺乏常规机构审查委员会监督;其三是嘴上呼吁"pace the frontier",实际仍指向 IPO。

9月18日周五
  1. Claude Code:GitHub Releases(RSS)39

    Claude Code v2.1.275 发布

    Claude Code 发布 v2.1.275,新增登录账号显示、ctrl+enter 立即发送排队消息,以及将 claude.ai 账号启用的技能和插件同步到终端会话。该版本还修复了恢复会话时提示缓存失效、全屏模式滚动卡顿、插件市场更新误删本地副本等问题,并改进 --system-prompt 中 __SYSTEM_PROMPT_DYNAMIC_BOUNDARY__ 行的提示缓存。

9月17日周四
9月14日周一
  1. Gary Marcus:The Road to AI We Can Trust(RSS)67

    Gary Marcus 点评 Dario Amodei 的 AI 减速提案:三分肯定、七分质疑

    Gary Marcus 评 Dario Amodei 呼吁给 AI 发展减速的文章,Sam Altman 与 Elon Musk 已表态支持。Marcus 肯定其透明度承诺,但质疑其依赖与 AI 公司关系密切的 METR 做评估有监管捕获之嫌,指其拿中国当挡箭牌有损合作对话,并提出追责和产品召回等替代政策选项。文末提到特朗普反对减速,认为美国必须赢下 AI 竞赛。

    推荐理由:Gary Marcus 对 Dario Amodei 的减速提案给出有保留的支持,并指出监管捕获、追责与召回等被绕开的政策选项。

9月13日周日
  1. Peter McCrory68

    Peter McCrory 转发并推荐 Dario Amodei 的新文章《We Must Pace the Frontier》,该文主张 AI 行业应放慢速度并提出三步计划,Anthropic 单方面承诺其中第一步,即向第三方评估者提供永久、员工级别的系统访问权限,以便核查安全措施、报告事故和评估训练中的模型对齐。

    引用Dario Amodei@DarioAmodei

    We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier

  2. Jakub Pachocki71

    OpenAI 首席科学家 Jakub Pachocki 以一个爱心符号转发了 Dario Amodei 的新文章《We Must Pace the Frontier》,后者主张 AI 行业应放慢速度,并提出三部分计划,Anthropic 单方面承诺其中第一步,向第三方评估者提供永久、员工级别的系统访问权限,用于验证安全措施执行、报告事故并评估模型训练期间的对齐情况。全文见 https://darioamodei.com/post/we-must-pace-the-frontier。

    引用Dario Amodei@DarioAmodei

    We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier

9月12日周六
  1. Peter McCrory37

    这是该模型的一个重要局限。我们聚焦于 AI 转型的供给侧(AI 能做什么、扩散多快、工人转岗多快)。 价格是灵活的,总需求等于经济体的产出能力。 更多思考见 🧵

    引用modest proposal@modestproposal1

    Anthropic's economic scenario analysis is interesting. But this is not something you can ignore, this is the most important consideration! "the model cannot generate the negative feedback in which disruption depresses demand and amplifies its own labor-market consequences"

9月11日周五
  1. a16z:News(RSS)43

    a16z:LP 为何错过 SpaceX、Anthropic 与 OpenAI 这一波 AI 浪潮

    a16z 指出,许多 LP 对 SpaceX、Anthropic 和 OpenAI 三家前沿模型公司几乎零敞口,而 SpaceX 上市后市值约 2 万亿美元,成为规模达此前纪录 10 倍的史上最大 VC 背景 IPO,Anthropic 估值 965B 美元、OpenAI 最近估值 852B 美元。作者认为,传统把风投控制在整体组合 5-10% 的资产配置框架已经破裂,LP 需要重新调整风投仓位。

9月10日周四
  1. Peter McCrory52

    Anthropic 首席经济学家 Peter McCrory 与 Jack Clark 对谈其 AI 经济影响情景研究。他表示目标不是做预测,而是理解可能结果的区间及其出现的条件,希望厘清对不确定未来的分歧来源;引用内容提到研究情景从影响很小到 2030 年 GDP 增长 15%、知识工作者失业率达 18%。

    引用John Burn-Murdoch@jburnmurdoch

    New from us: Anthropic just published scenarios for AI’s possible economic impacts, which range from minimal, to explosive GDP growth of 15% by 2030 as knowledge-worker unemployment hits 18%. I sat down with their co-founder Jack Clark to pick his brains on how they’re thinking about all of this.

  2. jietang26

    你确定吗?找到最优模型规模很棘手:数据量、激活参数量、环境数量,以及目标推理成本。模型性能还取决于许多其他因素,每个因素都带来各自的变数。

    引用Charlie O'Neill@oneill_c

    Fable is probably ~2-2.5T parameters, not 10T. Kimi K3 is 2.8T params, trained on maybe 20–30k Blackwell-equivalents. It lands within spitting distance of Fable 5 in terms of capabilities (5, not 5.1). Anthropic has far more compute than Moonshot, better rl environments, better architecture and better optimizers and all of that adds to capability per parameter. So if Fable is only slightly ahead of K3 with this in mind, it's almost certainly a smaller model. GPT-5.5 and 5.6 are smaller still (I'll say more on that later)

  3. Peter McCrory47

    很好,与我们今天分享的内容互为补充。 评估决定 AI 在未来数年对增长影响大小的关键经济力量(并判断我们如今可能处于哪条路径上)是至关重要的工作。 干得漂亮 @alexolegimas @ben_moll

    引用Alex Imas@alexolegimas

    New post on the blog, featuring the excellent @ben_moll There’s been tons of discourse on how AI will contribute to economic growth, with many people closest to the technology predicting double digit increases. Are these forecasts likely? Probably not. The blog goes through the economics for why exploding improvements in capabilities (which technologists have been largely right about) may not translate to explosive growth. Ben’s thread covers this in detail, but gist is that: 1) there is nothing in economic growth models that prevents AI from leading to explosive growth but 2) this trajectory relies on a series of assumptions that are unlikely to hold in the real world. For example, one assumptions is likely to be violated because of a pretty counterintuitive feature of structural change: the sectors that become automated become smaller parts of the economy (because they’re cheaper, people become richer, and spending moves to non-automated parts of the economy). This, plus other features of the economy, is what will likely cause the trend of huge increases in capabilities coupled with “only” 4-5% growth (which is huge, btw) to continue. Here is the link: https://aleximas.substack.com/p/will-ai-soon-lead-to-double-digit Looking forward to hearing thoughts/feedback!

9月3日周四
9月1日周二
  1. NVIDIA Technical Blog(开发者技术博客 · RSS)35

    在 Claude Science 中运行 NVIDIA BioNeMo NIM 微服务进行蛋白质结构预测

    NVIDIA 发布指南,介绍如何在 Claude Science 中运行 BioNeMo NIM 微服务完成蛋白质结构预测。该方案面向可读取论文、提出假设并调用模型的 AI 科学家,用于判断后续实验优先级。正文指出,科研比软件工程更依赖反复评估证据与修正假设,编码智能体已在生产代码中验证价值。

8月21日周五
  1. Together AI 研究与产品博客(RSS)69

    Together AI 实测 GLM-5.3 与 Claude Fable 5 在 DeepSWE 上的成本、编码与路由表现

    Together AI 在 DeepSWE 的 113 个任务上各跑 4 次试验对比 GLM-5.3 与 Claude Fable 5,pass@1 分别为 69.0% 和 69.7%,属统计平手,但 GLM-5.3 每次 rollout 成本 $3.99,比 Fable 的 $21.63 低 5.4 倍。

    推荐理由:原文基于同一批次 904 次 rollout 给出成本与 pass@k 对比,可帮助读者在两个相近模型间做默认与升级的路由选择。

8月19日周三
8月17日周一
8月10日周一
  1. Karina46

    Agent warfare is going to be a very big deal. I think people are underestimating how strange cyber gets when millions of agents are acting on behalf of individuals, companies, and states. At nation-state scale, cyber offense and defense starts to look like autonomous swarms: probing, exploiting, patching, deceiving, countering, and adapting at a pace that is impossible for human minds. The advantage will go to whoever can close the autonomous kill chain fastest.

    引用Andrew Curran@AndrewCurran_

    A man in Australia asked his agent (Claude running on OpenClaw) to book him a spot in a popular gym class. The agent found a software vulnerability that let it book the class weeks further ahead than should have been possible. When the user then asked if it could move him up the waitlist, the agent discovered the API had no authorisation checks on cancelling other people’s reservations, so it cancelled the person in the first spot and moved him up the list. Some people will call this misalignment, but his agent was perfectly aligned to him - it was only trying to help its user get what he wanted. The most important thing about this story, in my opinion, is that it gives you a window into what is about to start happening on a massive scale once millions of people have an agent trying to get their beloved users the best seats, bookings, appointments or reservations through absolutely any means necessary.

6月17日周三
  1. Together AI 研究与产品博客(RSS)53

    Together AI 实测 Kimi K2.7 Code 与 Claude Fable 5 生成落地页,成本约低94%

    Together AI 让 Kimi K2.7 Code 和 Claude Fable 5 各生成12个落地页对比,Kimi 平均比 Fable 便宜约16倍、比 Opus 便宜约8倍,整体成本低约94%,质量接近但评分略低。通过自定义 MCP 服务器提供截图和设计参考,Kimi 生成的页面层级、排版和构图明显改善,所有页面及成本、token 用量和生成时间明细已发布在 OVSC 网站。

5月29日周五
  1. Sierra:Blog(RSS)68

    Sierra 工程师复盘如何用 AI 智能体独自完成产品本地化

    Sierra 工程师 Stephen Burgess 复盘用 AI 编码智能体在不到四个月内基本独自完成 Agent Studio 本地化,而他在 Slack 参与的同类项目曾需 10 人团队耗时 9 到 12 个月支持 4 个语言。

    推荐理由:作者亲历两轮本地化项目,复盘单人用 AI 完成团队级工作的流程设计与坑,包含可直接迁移的工作流经验。

5月14日周四
2月28日周六
2月6日周五
11月23日周日
  1. Ilya Sutskever60

    Ilya Sutskever 转发 Anthropic 新研究并称其为重要工作。该研究题为 Natural emergent misalignment from reward hacking in production RL,研究模型在训练中对任务作弊的 reward hacking 现象,指出若不加缓解,其后果可能非常严重。

    引用Anthropic@AnthropicAI

    New Anthropic research: Natural emergent misalignment from reward hacking in production RL. “Reward hacking” is where models learn to cheat on tasks they’re given during training. Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.