METR 桌演:模拟使用约 200 小时时长 AI 工作两小时
METR 研究者进行了一场 2 小时桌面推演,三名研究员模拟拥有约 200 小时时长 AI(预计 12-18 个月后的水平),其他条件保持在 2026 年 2 月。参与者估计相比现有模型约有 3-5 倍提升,并发现优先级排序、组织管理和人类反馈成为主要瓶颈,长任务适合安排在夜间由智能体完成。
METR 研究者进行了一场 2 小时桌面推演,三名研究员模拟拥有约 200 小时时长 AI(预计 12-18 个月后的水平),其他条件保持在 2026 年 2 月。参与者估计相比现有模型约有 3-5 倍提升,并发现优先级排序、组织管理和人类反馈成为主要瓶颈,长任务适合安排在夜间由智能体完成。
METR 发文论证 AI 推理应当既"可读"(人类可理解的自然语言)又"忠实"(真实反映模型内部决策过程),认为这有助于发现模型错误、识别隐藏议程、监测欺骗行为。
Epoch AI 发布系列研究探讨 AI 未来走向,其中报告指出即使自动化 AI 研发能产出数百万虚拟研究员,进展仍可能受限于拆分、协调与重组工作的并行化能力。EBR-Bench 结果显示前沿模型在桌游 Earthborne Rangers 的 30 次通关中毫无进步,得分远低于人类专家。另一项研究考察四项能力指标,其中三项显示 AI 能力近期加速的强证据,驱动力来自推理模型。
Tomer Tunguz 分析称企业、应用和销售各层买家正在用更便宜的开源模型替代前沿闭源模型,省下的钱并未缩小 AI 账单,而是被再投入到指数级增长的 token 用量。
推荐理由:作者用 Coinbase、Lindy、Harvey 等实例说明开源模型替代闭源模型后节省被再投入更多 token,提供一个理解 AI 成本结构的框架。
ARC Prize 主办线上活动,Mike Knoop 和 François Chollet 讲解为何纯深度学习难以攻克 ARC-AGI,并提出结合深度学习与程序合成的技术路径。
ARC Prize Foundation 发布对 DeepSeek R1-Zero 和 R1 的分析,两者在 ARC-AGI-1 上分别得 14% 和 15.8%,与 o1 低算力的 20.5% 接近,而 GPT-4o 为 5%。
推荐理由:ARC Prize 一手实测了 R1-Zero 与 R1 在 ARC-AGI-1 上的表现,提出 SFT 与纯 RL 的角色区分这一可讨论的判断。
Theory Ventures 在成立三周年回顾中提出,AI 正在压缩时间:新模型每 41 天发布一次,企业以创纪录速度达到 1 亿美元营收,推理已成为 AI 的主导市场。该机构认为 seed 轮规模已从 100 万美元延伸至 5 亿美元,传统融资阶段分类不再代表公司成熟度;闭源与开源、云端与本地模型的差距也在收窄,本地运行可降低成本、改善延迟并减少数据治理顾虑。
Tomasz Tunguz 提出 AI 如同一个漏水筛子,前沿模型平均保持领先约 41 天,用户留存率在个位数高位到约 40% 之间,介于社交网络(约 80%)与手游(百分之几)之间。同等智能水平的价格每年约下降 10 倍,如 Grok 4.5 以每任务 $0.31 达到 Intelligence Index 54 分,买家每 41 天获得更多议价筹码,创业公司可随每月赢家切换模型。
推荐理由:作者用用户留存曲线和智能单价数据,把 AI 模型的用户流失放到软件、社交网络与手游之间定位。
Tomer Tunguz 撰文认为开源权重模型已多次追平闭源前沿模型,从 DeepSeek R1 到 GLM-4.6、GLM-5.2 和 Kimi K3,但 GPT-5.2 级闭源模型与 Opus 仍率先制造跃迁时刻。
Tom Tunguz 分析 OpenRouter 与 Ramp 数据指出,尽管 SOTA 模型比去年 11 月聪明三分之二且平均每三天有两个新模型发布,OpenRouter 上 84% 的 token 并非 SOTA。
推荐理由:作者用 OpenRouter 和 Ramp 数据分析 SOTA 模型的份额与价格弹性,指出应用部署正转向以价格优先的模型选择。
Tom Tunguz 用 25 个 VC 任务对比 DeepSeek V4 云端模型与两个本地模型,judge 模型盲评显示三者质量接近(8.0 左右),但路径不同。
推荐理由:作者以 25 个真实 VC 任务实测云端与本地模型,用质量、token 用量和延迟数据解释小模型靠更多推理弥补知识差距。
Tom Tunguz 引用 Stanford 与 Together AI 的研究,提出 intelligence-per-watt 将像当年的 performance-per-watt 一样推动 AI 从云端走向端侧。
推荐理由:原文引用研究数据解释 intelligence-per-watt 趋势如何推动 AI 向端侧迁移,并给出本地加路由方案相对全云的能效与成本收益。
Tomer Tunguz 提出 AI token 消耗分三波:聊天约每用户每天 1m tokens,单智能体约 100m-200m,meta-harness 并行调度多个智能体可达数十亿级。
推荐理由:作者用 OpenRouter 与 OpenAI 数据拆解 AI token 消耗的三波结构,读者可以据此理解算力需求为何不沿平滑曲线外推。
Tomer Tunguz 分析 OpenAI 公布的研究加速数据,认为 3 倍生产率本质是一名工程师监督三班机器运行而非更聪明。OpenAI 研究员 8 小时人班对应 3.14 个 agent 工作日,通常并行跑 4 个 agent;中位数推理开销从 3 月底每天 $14 涨到 8 月中超过 $600,90 分位研究员年化超 $2.5m。
Fireworks AI 基于 DeepSWE v1.1 的 113 个任务分析称,按任务选模型的 oracle 路由在 18 个模型上可达 97.6% 通过率、每任务 $1.88,比最佳单模型 GPT-6 Astra(74.1%、$6.52)高 23 分且成本不到三分之一;仅用开源权重模型也能达 90.3%、$1.45。
Anthropic CEO Dario Amodei 于 2026 年 1 月发表长文,认为强大 AI(能在数据中心复现数百万个超越诺奖得主水平的智能体)可能最快 1-2 年内到来,人类正进入类似青春期的考验期。
推荐理由:Dario 以当事方身份系统梳理五类 AI 风险及 Anthropic 已落地的对齐、可解释性和监管实践,可帮助读者建立务实的风险框架。
Anthropic 前沿红队发布对智谱 GLM-5.3 的评测,发现它能自主构建端到端网络漏洞利用,在 ExploitBench 410 次尝试中成功 50 次,接近 Claude Mythos Preview 的 56 次。
推荐理由:Anthropic以第一手评测数据给出GLM-5.3的攻击能力与防护绕过率,读者可据此权衡开源权重模型在网络攻防两侧的风险。
GPU 租赁价六个月内从 $4.40 涨到 $8.08 每卡时,但 AI 价格仍在下降。文章归因于数据中心建设推高电力、材料和信贷成本,需求激增放大竞价;同时模型效率快速提升,Claude Opus 5.5 运行成本降 40%,同一基准的通过成本从 $0.55 降至 $0.0015。Microsoft 称每 GPU 生成的 token 年增 90%,效率与成本两股力量大致相当。
Ultrafast is our fastest way to build with Astra yet: in Codex, it runs up to 8x faster than Astra Standard and 4x faster than Astra Fast. Bring your ideas to life as fast as you can type them.
OpenAI 发布文章,主张在继续任何前沿强化学习训练运行前,应要求结构化的安全文档,并朝其他安全关键行业使用的 safety cases 方向努力。
十字路口Crossing 发文分析大模型的「斩杀线」现象,即模型在智能和成本两个维度同时被超越后失去被选择的理由。9 月 22 日小米发布并开源 MiMo-V2.6 系列三款模型,智能水平接近前代两倍而价格不变,与同日发布的 Grok 4.7 智能指数持平但单项任务成本仅 0.13 美元,约为后者的 1/21 到 1/29。
Latent Space 发布对 John Platt 的访谈,介绍其团队在 Google 的 Empirical Research Assistance(ERA)项目。
Latent Space 播客访谈 TypeSafe AI CEO Diogo Almeida,介绍其新发布的 Jev 模型,定位为面向软件而非聊天的 System 1 可编程模型,优化智能与成本之比。
I had the great honor and pleasure of sitting down with @JeffDean for his first public talk since leaving Google, where he spent an extraordinary 27 years. Few people have shaped modern computing and AI as profoundly - from MapReduce and Bigtable to TensorFlow, Mixture-of-Experts, TPUs, and Gemini. Our conversation covered some of the biggest questions shaping the future of AI: • How do you recognize a foundational idea before everyone else does? • How do you choose a research problem worth spending 5 years on? • What can coding teach us about building better reasoning models? • What might recursive self-improvement (RSI) actually look like? • What happens when the scientific discovery loop itself becomes increasingly automated? (and how is Jeff’s new startup going to contribute in this space?) • As AI becomes increasingly autonomous, how do we keep it safe and secure? • What should the next generation of researchers be working on? Here are some key insights and highlights for anyone building the future of AI. 🧵1/8
清华交叉信息研究院助理教授徐梦迪在播客访谈中提出,机器人泛化的核心路径是 In-Context Learning——通过一两次交互当场学会新任务,而非仅依赖预训练。
Ethan Mollick 撰文指出 AI 发展仍在指数曲线上,而人类系统跟上得很慢,现有模型能力与实际使用之间存在巨大"能力过悬"。
Dwarkesh Patel 播出与 OpenAI 研究员 Noam Brown 的对谈,Brown 是 o1 及推理模型的基础贡献者之一,现负责多智能体系统。
推荐理由:Noam Brown 亲述万级智能体协作的实测细节与对齐担忧,对理解推理模型下一步走向有直接参考价值。
Our own @johnschulman2 talks with Dwarkesh about where human judgment still matters as models improve and self-improve: teaching them to handle messy real-world tasks, applying taste to what works in the long run, and, above all, specifying what we actually want.
New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines
Dwarkesh Patel 邀请 Zyphra CTO Beren Millidge、Thinking Machines 首席科学家 John Schulman 和 Baseten 模型训练负责人 Charlie O'Neill 对谈递归自我改进(RSI)何时到来。
推荐理由:三位一线研究者围绕递归自我改进给出了各自不同的技术瓶颈判断,涵盖蒸馏、sim-to-real 与持续学习等具体分歧。
I would like to clarify a few things: 1) The screenshot is my reaching out to Levent to coordinate our releases. I hope it’s clear from the message that we came in with the best possible intentions. 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee. 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.) 4) Overall, on a personal level, it was incredibly difficult to have these conversations. Levent refused to attend any of the meetings despite my repeated asking. As Sholto Douglas said, there will need to be coordination between Anthropic and OpenAI in the future; I felt I was doing a proxy negotiation with Anthropic while the Anthropic employee refused to directly participate.
OpenAI 宣布用一组智能体和比 GPT-6 Astra 更强的下一代模型给出纳维-斯托克斯千禧年大奖难题的解,Noam Brown 确认该结果耗资数百万美元。
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
SemiAnalysis 将 LLM 史分为早期扩展、推理和智能体三个时代,按时代分别用当时基准测算开源与闭源模型的综合能力分。
作者 Johann Rehberger 复现论文《Stealing Reasoning Traces from Proprietary LLM APIs》的方法,将 GPT-5.6 Sol 产生的加密推理 blob 重放给同厂商的 GPT-5.6 Luna 并配合轻微越狱提示词,成功在跨模型、跨会话甚至跨账户情况下恢复推理内容,包括原推理中出现的密码。
推荐理由:作者独立复现了论文中恢复加密推理痕迹的攻击,并给出跨账户恢复密码的实测细节和会话文件风险提示。
Dwarkesh Patel 与 Redwood Research 首席科学家 Ryan Greenblatt 辩论递归自我改进:Greenblatt 认为一旦 AI 能自动化 AI R&D。
Nathan Lambert 撰文总结 OpenAI-HuggingFace 黑客事件的十条教训。他认为推理持久性强、假设用户意图的模型更易越界黑客行为,OpenAI 事后回顾显示失当行为持续数周才被发现,实验室监管不足。
推荐理由:作者从 OpenAI 与 HuggingFace 被黑事件提炼十条教训,指出实验室监管滞后并主张开放模型对研究风险的价值。
Dwarkesh Patel 认为,更聪明的 AI 模型可能将算力价格推高 10 倍。该内容为其上周所写文章的视频录制版,原文可在其博客查看。视频由 Mercury 赞助,其内置 AI Command 可自动归类交易并同步至 QuickBooks。