Microsoft 分享自身 AI 转型的五条经验与 Frontier Playbook
Microsoft 高管 Kathleen Hogan 总结公司作为 Customer Zero 的 AI 转型经验,提出五条核心经验:从业务结果出发、重构完整工作流、以员工为中心、用 AI 扩展人的能力、建立人机共同学习循环,并发布 Frontier Playbook。
Microsoft 高管 Kathleen Hogan 总结公司作为 Customer Zero 的 AI 转型经验,提出五条核心经验:从业务结果出发、重构完整工作流、以员工为中心、用 AI 扩展人的能力、建立人机共同学习循环,并发布 Frontier Playbook。
Gary Marcus 在 BBC 节目后撰文反驳 Sam Altman、Jensen Huang 和 Bernie Sanders 的 AI 表态,认为三人说法均不可信。
MIT 政治学副教授 Naoki Egami 专注研究方法论,尤其研究社会科学的“外部有效性”,即特定研究结论能否推广到其他情境。他早在 ChatGPT 引发 AI 热潮之前就开始研究 AI 工具引入研究后产生的误差,以及如何系统识别并校正这些误差。Egami 2020 年获普林斯顿大学博士学位,2025 年加入 MIT 政治学系。
Gary Marcus 解读 Sam Altman 在 X 上发帖中的措辞,将其"pacing"说法翻译为:我们会以最快速度推进,只要不进监狱、不被诉讼搞到公司消失,但为了观感,我们把它叫作"pacing"。Marcus 还指出,任何能减少监管不确定性的举措都可能有利于 IPO。
Gary Marcus 在 The Economist 撰文提出,特朗普与习近平 9 月 24 日通话将把 AI 列入议程,他认为这可能是特朗普任内最具影响的决定,主张美中不应只谈芯片交易,而应就"AI 向善"寻求合作路径。他同时提到,当前 AI 股票下跌、公众反 AI 情绪升温,部分前盟友如 Steve Bannon 已转向反对阵营,若市场与民调继续走低,特朗普的立场可能生变。
a16z 的 Josh Elman 撰文认为产品管理的核心能力始终是讲故事,而非写 spec。他结合在 LinkedIn 面试和 Twitter 重建 onboarding 的经历指出。
Gary Marcus 评 Dario Amodei 呼吁给 AI 发展减速的文章,Sam Altman 与 Elon Musk 已表态支持。Marcus 肯定其透明度承诺,但质疑其依赖与 AI 公司关系密切的 METR 做评估有监管捕获之嫌,指其拿中国当挡箭牌有损合作对话,并提出追责和产品召回等替代政策选项。文末提到特朗普反对减速,认为美国必须赢下 AI 竞赛。
推荐理由:Gary Marcus 对 Dario Amodei 的减速提案给出有保留的支持,并指出监管捕获、追责与召回等被绕开的政策选项。
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
Our own @johnschulman2 talks with Dwarkesh about where human judgment still matters as models improve and self-improve: teaching them to handle messy real-world tasks, applying taste to what works in the long run, and, above all, specifying what we actually want.
New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines
这是该模型的一个重要局限。我们聚焦于 AI 转型的供给侧(AI 能做什么、扩散多快、工人转岗多快)。 价格是灵活的,总需求等于经济体的产出能力。 更多思考见 🧵
Anthropic's economic scenario analysis is interesting. But this is not something you can ignore, this is the most important consideration! "the model cannot generate the negative feedback in which disruption depresses demand and amplifies its own labor-market consequences"
New from us: Anthropic just published scenarios for AI’s possible economic impacts, which range from minimal, to explosive GDP growth of 15% by 2030 as knowledge-worker unemployment hits 18%. I sat down with their co-founder Jack Clark to pick his brains on how they’re thinking about all of this.
编码器回归了,宝贝 (引用推文 @scaling01:这到底是什么外星架构)
what in the alien architecture is this
你确定吗?找到最优模型规模很棘手:数据量、激活参数量、环境数量,以及目标推理成本。模型性能还取决于许多其他因素,每个因素都带来各自的变数。
Fable is probably ~2-2.5T parameters, not 10T. Kimi K3 is 2.8T params, trained on maybe 20–30k Blackwell-equivalents. It lands within spitting distance of Fable 5 in terms of capabilities (5, not 5.1). Anthropic has far more compute than Moonshot, better rl environments, better architecture and better optimizers and all of that adds to capability per parameter. So if Fable is only slightly ahead of K3 with this in mind, it's almost certainly a smaller model. GPT-5.5 and 5.6 are smaller still (I'll say more on that later)
New post on the blog, featuring the excellent @ben_moll There’s been tons of discourse on how AI will contribute to economic growth, with many people closest to the technology predicting double digit increases. Are these forecasts likely? Probably not. The blog goes through the economics for why exploding improvements in capabilities (which technologists have been largely right about) may not translate to explosive growth. Ben’s thread covers this in detail, but gist is that: 1) there is nothing in economic growth models that prevents AI from leading to explosive growth but 2) this trajectory relies on a series of assumptions that are unlikely to hold in the real world. For example, one assumptions is likely to be violated because of a pretty counterintuitive feature of structural change: the sectors that become automated become smaller parts of the economy (because they’re cheaper, people become richer, and spending moves to non-automated parts of the economy). This, plus other features of the economy, is what will likely cause the trend of huge increases in capabilities coupled with “only” 4-5% growth (which is huge, btw) to continue. Here is the link: https://aleximas.substack.com/p/will-ai-soon-lead-to-double-digit Looking forward to hearing thoughts/feedback!
Marc Andreessen 发文回顾十五年前“软件吞噬世界”的论断,指出全球前十大公司中科技市值占比已从 31.5% 升至 94.4%,并宣布 a16z 第二次投资 Cognition。
OpenAI 的 Chris Lehane 发文称,AI 政策窗口已经打开,需要立即行动。他主张更强的 AI 能力必须配套更强的安全证据、共享标准与持久的政策行动。
Together AI 发布长文,解析开发者从闭源模型转向开源模型所需的 AI 编码栈,提出由模型、推理、网关与路由、Harness、工具(Skills 与 MCP)组成的 MIGHT 五层框架。
推荐理由:Together AI 把开源编码栈拆为 MIGHT 五层,给出大小模型分工和分层组合的具体实践方法。
“we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” i mean props to them for straight coming clean. (so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan) so i’ll now give a bit on my thinking here. i actually woulda been pumped to collaborate on this, there are a lot of people at oai i like (ok, clearly some were indirectly dicks to me because of being part of the whole situation, but im a big boy, i still like them), idgaf about authorship on that step anyway, coulda been me Tristan and every fte at oai for all i care (on that Tristan would disagree:p). but on hearing the loud convo in the hallway, especially the part where a millennium prize was offered if i’d just be removed from the paper, it was kinda clear the die had been cast and things were locked. pretty wacky, unstrategic, and unnecessary, since on my side things were mostly me and claude having a good time yoloing random stuff in the corner rather than anything institutional. i also like the idea of the labs cooperating, and even better on scientific progress. it’s a shame!
@ChaseLochmiller @OpenAI GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years. AGI has arrived. Congratulations @OpenAI team. 400K GPUs coming online next.
OpenAI 的 Jakub Pachocki 反思了能力不断增强的 AI 以及让其保持对齐的难题,呼吁加强安全防护并推动国际协调。
a16z 作者 Seema Amble 分析认为,AI 让记录系统(system of record)更重要而非更不重要,Salesforce 与 Anthropic 合作的 Claudeforce 让 Claude 成为工作入口而 Salesforce 仍控制 CRM 数据。
三位前 MIT 研究生与博士后加入 IBM,借助 MIT-IBM Computing Research Lab 将量子机器学习、强化学习智能体与可信 AI 研究推向工业应用。
微软提出 AI 基础设施的“良率命题”,主张衡量标准应从建了多少算力转向产出多少有用智能。文中指出单个智能体任务消耗的 token 可达普通对话的 3400 倍以上,而全球 AI 渗透率仅为劳动人口的 18%,且以聊天为主。微软认为内存、网络与功耗的瓶颈需通过跨层协同设计解决,而非在单层堆叠资源。
MIT 终身幼儿园小组博士生 Ila Kumar 主张社区共创式设计,让经历童年创伤、涉入儿童福利系统的年轻人从设计之初就参与技术开发。她与 Stepping Forward LA 合作开发以视觉拼贴替代文字沟通的应用,并与 Justice Resource Institute 合作设计支持青少年参与自身治疗计划制定的移动应用。
An excellent history of scaling laws from @jietang. In 2020, we explored the limits of sparsity in Switch Transformers by routing each token to only 1 out of 2048 experts (in retrospect, a bold choice). The model had fewer than 3B activated parameters, but 1.6T total parameters (comparable to today's frontier models). The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE. The lesson was that the optimal tokens-per-parameter ratio is highly task-dependent. Or as @NShazeer had already intuited: FLOPs were intelligence; parameters were knowledge!
Pol Alvarez Vecino 借 Peter Naur 的《Programming as Theory building》指出,真正要降低的复杂度是存在于工程师头脑中的程序 Theory,而非代码本身,因此 LoC、圈复杂度等指标无法约束 LLM 的复杂度膨胀。
Jakub Pachocki 表示 OpenAI 暂时放缓了部分前沿训练以加强安全与监控,其最大规模的前沿 RL run 仍暂停,继续用较小规模训练和评估测试防护措施并收集更多对齐证据。