METR 研究员 Thomas Kwa 撰文澄清 AI 时间跨度研究的局限与核心结论
METR 时间跨度论文主要作者 Thomas Kwa 撰文澄清对该研究的常见误读,指出时间跨度是指以 50% 成功率可替代的串行人类劳动时长,而非 AI 可独立工作的时间,且测量误差较大、跨领域差异可达数量级,如视觉计算机使用任务低 40-100x。
METR 时间跨度论文主要作者 Thomas Kwa 撰文澄清对该研究的常见误读,指出时间跨度是指以 50% 成功率可替代的串行人类劳动时长,而非 AI 可独立工作的时间,且测量误差较大、跨领域差异可达数量级,如视觉计算机使用任务低 40-100x。
斯坦福 HAI 发文梳理 AI 加速科学发现的进展:Brian Hie 团队打造的 Evo 2 是迄今最大的生物学 DNA 语言模型,基于 9 万亿碱基对、400 亿参数训练,可像聊天机器人一样补全基因序列。Emma Lundberg 团队则正构建模拟人类细胞的基础模型,用于加速药物发现与个性化医疗,并为此开发了 Biomni 供生物学家交互分析数据。
Tomer Tunguz 撰文分析转售推理的公司如何保持 30 点以上毛利,指出成本加成定价会随推理商品化而压缩至零,客户会绕开加价直接对接原始 API。
Tomer Tunguz 分析称 Anthropic 的算力支出约为其薪酬总额的 2.3 倍,约 5000 名员工对应 2026 年约 $10b 推理与训练支出,即每人每年约 $2m;而软件行业前 1% 公司每工程师每年 AI 支出为 $89k,中位数仅 $137。
VC Tomer Tunguz 撰文分析 AI 时代企业数据外流风险,引用 Satya Nadella 的“反向信息悖论”与 Palantir CEO Alex Karp 关于前沿实验室“窃取业务权重与 alpha”的言论。
Tomer Tunguz 分析,超大规模云厂商在 2026 年 Q2 财报电话会上称 AI 产能持续受限,HBM3e 内存涨价 20%、HBM4 预计翻倍。
推荐理由:作者用各实验室定价与算力供给数据,解释分层定价如何维持杰文斯悖论,并指出路由层可能成为新的战略环节。
Tom Tunguz 分析 OpenRouter 与 Ramp 数据指出,尽管 SOTA 模型比去年 11 月聪明三分之二且平均每三天有两个新模型发布,OpenRouter 上 84% 的 token 并非 SOTA。
推荐理由:作者用 OpenRouter 和 Ramp 数据分析 SOTA 模型的份额与价格弹性,指出应用部署正转向以价格优先的模型选择。
测试时训练让模型在回答提示词时就地做梯度更新,权重随使用而改变,从而把不断增长的 KV-cache 折叠为固定大小的权重,内存占用不再随上下文线性增长。斯坦福针对小模型的研究显示其推理速度最高可提升 2.7 倍,In-Place TTT 还能免重训把 4b 模型提升到有竞争力的 128k 上下文表现。
Eugene Yan 撰文分析复杂度偏见为何存在,指出复杂方案通过信号化努力、掌控力、创新和功能多样性而更受论文评审和晋升机制青睐,而简单方案更易采用、构建、扩展和运维。文章举例说明简单机器学习方法常不逊于复杂方法,如点积在推荐检索上优于神经协同过滤、树模型在 45 个中型表格数据集上超过深度神经网络,并建议聚焦问题复杂度而非方案复杂度,善用奥卡姆剃刀。
Eugene Yan 回顾 2024 年,将 2023 年的原型规模化落地为服务客户的 ML/LLM 生产系统,并发布 6 篇长文,其中与友人合写的《What We've Learned From a Year of Building with LLMs》还由 O'Reilly 出版成书。
Andrej Karpathy 的 Google 暑期实习工作最终成为 CVPR 2014 Oral 论文《Large-scale Video Classification with Convolutional Neural Networks》。
一批成立不满一年的数据公司估值急速膨胀,诞生了25亿美金的 UniPat 和多家3-8亿美金估值的公司,被作者概括为给模型制作练习题的生意,即采集专家解题轨迹或搭建 RL environment 交给模厂做专项训练。
HPE 提出企业 AI 从按 token 消费转向自建容量的判断框架:当需求稳定、可预测且规模足够时,拥有算力可能比逐次购买更经济。文中引用 Deloitte 2026 企业 AI 状况报告称,2025 年员工 AI 使用率上升 5%,至少 40% 的 AI 项目进入生产的公司比例预计半年内翻倍。企业需先回答三个问题:需求是否稳定、在什么使用水平下自建更划算、能否通过采用与治理让容量保持高产。
这是几个月前我还在 Google 时录制的,真的是一次非常有趣的对话!
How does this only have 21,000 views in 8 days? Chat with @JeffDean (then Google) and Bill Jia about large scale AI models. https://youtu.be/BVQSWeK2Nrw?si=MvMXdAEqjQE4HBOb
在 AI x Science 领域,生物技术公司正沿两条路径适应 AI:Foundries 用新一代测序、高通量显微和物理自动化将实验数据生成速度提升一个数量级,AI 让数据可读可预测,但差异化资产仍是实验数据本身;Navigators 则把 AI 嵌入公司日常流程,驱动更优决策与更快流程,无需专有模型或大规模数据集。
Latent Space 播客访谈 TypeSafe AI CEO Diogo Almeida,介绍其新发布的 Jev 模型,定位为面向软件而非聊天的 System 1 可编程模型,优化智能与成本之比。
峰瑞资本李丰撰文分析,认为2026年三季度全球流动性接近见顶,美元主导的资本市场进入存量博弈尾部,AI产业周期进入后半段。文章回顾2020年天量流动性如何催生本轮AI热潮,列举科技巨头资本开支转折的五个信号(如Alphabet二季度自由现金流转负59亿美元),提出投资重心应从讲大故事转向能靠AI赚钱的方向,如AI+应用、生物医疗与AI交叉及SaaS的AI化。
MIT 政治学副教授 Naoki Egami 专注研究方法论,尤其研究社会科学的“外部有效性”,即特定研究结论能否推广到其他情境。他早在 ChatGPT 引发 AI 热潮之前就开始研究 AI 工具引入研究后产生的误差,以及如何系统识别并校正这些误差。Egami 2020 年获普林斯顿大学博士学位,2025 年加入 MIT 政治学系。
New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines
Dwarkesh Patel 邀请 Zyphra CTO Beren Millidge、Thinking Machines 首席科学家 John Schulman 和 Baseten 模型训练负责人 Charlie O'Neill 对谈递归自我改进(RSI)何时到来。
推荐理由:三位一线研究者围绕递归自我改进给出了各自不同的技术瓶颈判断,涵盖蒸馏、sim-to-real 与持续学习等具体分歧。
编码器回归了,宝贝 (引用推文 @scaling01:这到底是什么外星架构)
what in the alien architecture is this
“we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” i mean props to them for straight coming clean. (so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan) so i’ll now give a bit on my thinking here. i actually woulda been pumped to collaborate on this, there are a lot of people at oai i like (ok, clearly some were indirectly dicks to me because of being part of the whole situation, but im a big boy, i still like them), idgaf about authorship on that step anyway, coulda been me Tristan and every fte at oai for all i care (on that Tristan would disagree:p). but on hearing the loud convo in the hallway, especially the part where a millennium prize was offered if i’d just be removed from the paper, it was kinda clear the die had been cast and things were locked. pretty wacky, unstrategic, and unnecessary, since on my side things were mostly me and claude having a good time yoloing random stuff in the corner rather than anything institutional. i also like the idea of the labs cooperating, and even better on scientific progress. it’s a shame!
We spent >20B tokens throwing @openai's Astra at every AI Engineering task we could think of, beyond cute Blender demos and fun games. https://latent.space/p/astra Here's everything Astra can do, and do so at <$6 an hour (serious): - choose and train models - label data (both helping you label and then using your labels for active learning) - keep pipelines saturated - instrument and read logs - deploy and debug entire systems in one shot - fan out and command and eval subagents (including agents running other models) - keep coherence over billions of tokens of a single agent thread. more to come on @swyx's coverage of the Fable- and Astra-class of 2026!
Regret the tone of my post on data centers yesterday. What I should have said: There were reasonable concerns about data centers 18ish months ago: water, taxes, jobs, electricity prices, the environment and what they would do to small towns. Well-structured data center projects have largely addressed these concerns today and we should be celebrating this. On balance, data centers are awesome for America in every way. On water: U.S. data centers use a fraction of what golf courses use. A lot of the numbers from 18 months ago were off by over 1000x. Newer data centers use closed-loop systems or recycled water. Should be required by every town approving a data center project. On taxes: looking only at sales-tax exemptions, as Ronan Farrow did, is the wrong way to evaluate this. Data centers pay significant property taxes. Loudoun County, which is the wealthiest county in America, now collects on the order of $1 billion a year from data centers. In Quincy, WA, data centers are more than half the property-tax roll. Over time, property taxes can go to zero while government spending increases in these towns. On jobs: this has been unambiguously awesome for blue collar Americans. Demand for electricians, plumbers, welders, HVAC techs, and contractors has gone vertical, and it is not a one-time construction job. These buildings get upgraded and expanded over time. That is why the building trades are fighting for them, and why some unions are now treating opposition to data centers as a reason not to endorse politicians. On power: the original fear was that households would pay for the incremental electricity demand in the form of higher prices. That is why the ratepayer-protection deals and the new large-load tariffs exist. The right structure is: the data center brings or pays for new generation and signs a contract long enough that existing customers are protected. Where that is happening, utilities are cutting or freezing residential rates and saying so on the record. Where it is not, people are right to object. Electricity prices are going down *today* in a number of large states because of data centers. On the environment: data centers overwhelming use natural gas today, which is the cleanest power source outside of nuclear, solar and wind. And the companies that are building the data centers are committed to carbon neutrality such that an equivalent amount of solar will likely be built. Maybe more importantly, the data centers need batteries to function effectively and these batteries can also sell energy back into the grid (which recently prevented blackouts in Texas). Over time, data centers will run on solar plus batteries. On the towns: Poverty in Quincy, WA fell from 29% to 6%. Data center taxes paid for a new high school, a hospital, a library, police and fire stations. This is happening in many left for dead former mill and farm towns that had no other bidder for the land. Data centers are actually reindustrializing parts of America and creating the kind of working-class jobs both parties have spent decades claiming to support. That should not be a partisan issue. Data centers can and should be awesome for America and they increasingly, overwhelmingly are. Supporting the outsourcing of data centers to China will likely age just as well as support for the outsourcing of high quality, blue collar manufacturing jobs to China has aged. When the facts change, I change my mind. I hope that reasonable people who had good faith reasons to oppose data centers at least consider updating their beliefs given the change in the facts over the last 18 months. This really matters for America. I will say I also think the idea of making data centers beautiful is a good one that has yet to be implemented. Data centers should be just as beautiful as Grand Central Station. We can learn a lot from the railroad buildout. Neoclassical revival ftw. Might write up open-weight AI tomorrow as this is equally essential to America.
Pyromind 创始人兼 CEO Kevin Ding 在播客中提出,AI 终局更像 Agent 蜂群而非超级基础模型一统天下,公司押注 AutoRL 而非仅做 RL as a Service。
🚢 Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow. Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.
SemiAnalysis 将 LLM 史分为早期扩展、推理和智能体三个时代,按时代分别用当时基准测算开源与闭源模型的综合能力分。
An excellent history of scaling laws from @jietang. In 2020, we explored the limits of sparsity in Switch Transformers by routing each token to only 1 out of 2048 experts (in retrospect, a bold choice). The model had fewer than 3B activated parameters, but 1.6T total parameters (comparable to today's frontier models). The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE. The lesson was that the optimal tokens-per-parameter ratio is highly task-dependent. Or as @NShazeer had already intuited: FLOPs were intelligence; parameters were knowledge!
Synopsys 的 Ravi Subramanian 在 DAC 2026 音频访谈中讨论了芯片设计中的物理学与 EDA 工具,重点谈及 3DIC 与热管理。他指出典型移动 SoC 约 2 到 25 亿门,而典型汽车 ECU 芯片约 70 亿门,功耗已直接影响到电动车续航。随着芯片变大,机械应力等原本的二三阶效应正变成一阶效应,签核需同时考虑机械与电学性能。
Dwarkesh Patel 认为,更聪明的 AI 模型可能将算力价格推高 10 倍。该内容为其上周所写文章的视频录制版,原文可在其博客查看。视频由 Mercury 赞助,其内置 AI Command 可自动归类交易并同步至 QuickBooks。
Google Cloud 在 Next 2026 上推出 Agentic Data Cloud,将数据、AI 模型与运营数据库统一为单一 System of Action。
Moonshot AI 于 7 月 16 日发布旗舰模型 Kimi K3,为 2.8T 参数 MoE 模型,权重定于 7 月 27 日开放,在 Vals AI 指数排第 2、Artificial Analysis 智能指数排第 3、Frontend Code Arena 排第 1。
推荐理由:作者基于榜单和架构细节分析 K3 对开源模型经济与中美竞争格局的影响,提供了可参考的行业判断框架。
英灵殿创始人 Odin 在播客中讲述自己高中辍学、自学考入浙大、离开 David Baker 诺奖实验室后投身 AI for Science 创业,已融资数千万美元。他提出要做「全模态分子世界模型」和「通用科学人工智能」,目标是造一台「新时代科学发现的蒸汽机」,把人类科学进程压缩上百年。团队约 30 人,曾让一名本科生带做 8 个靶点,每轮都做出 sub 纳摩尔级活性分子。
UC Berkeley 的 Aditya G. Parameswaran 等人提出,推理成本正快速趋近于零——GPT-4 级能力从 2023 年初的约 $30/百万 token 降至如今不到 $1,部分厂商已低于 $0.10,基准测试中推理价格年降幅在 9x 到 900x 之间,中位数约 50x。
Dwarkesh Patel 公布其 AI 大问题征文比赛的三篇获奖作品,比赛共收到超 600 篇投稿。
Dwarkesh Patel 撰文提出,实验室押注 RLVR 在数千环境中训练可通向 AGI,但仅在会话内学习不够,真正的下一个范式是让 AI 从部署后的真实经验中持续学习并回写到权重。
推荐理由:作者用计算机使用进度慢的例子说明可验证性之外还需可反复模拟性,并给出把部署经验蒸馏回权重的可能路径。
Dwarkesh Patel 撰文指出,过去几年 AI 进步主要来自更多更好的数据与算力,而非训练样本效率的提升;RL 可视为一种合成数据生成,需要大量高度定制的人类专家轨迹和 rubric。
推荐理由:作者用人类与模型的 token 消耗对比和 Chinchilla 常数推算,解释为什么数据而非训练技巧才是追赶前沿的关键。