跳到正文

#大佬观点

今日 3 条
9月17日周四
9月16日周三
  1. MIT News(RSS)23

    MIT 政治学者 Naoki Egami 如何用统计方法研究社会测量与 AI 工具误差

    MIT 政治学副教授 Naoki Egami 专注研究方法论,尤其研究社会科学的“外部有效性”,即特定研究结论能否推广到其他情境。他早在 ChatGPT 引发 AI 热潮之前就开始研究 AI 工具引入研究后产生的误差,以及如何系统识别并校正这些误差。Egami 2020 年获普林斯顿大学博士学位,2025 年加入 MIT 政治学系。

9月15日周二
  1. Gary Marcus:The Road to AI We Can Trust(RSS)30

    Gary Marcus 解读 Sam Altman 的"pacing"表态

    Gary Marcus 解读 Sam Altman 在 X 上发帖中的措辞,将其"pacing"说法翻译为:我们会以最快速度推进,只要不进监狱、不被诉讼搞到公司消失,但为了观感,我们把它叫作"pacing"。Marcus 还指出,任何能减少监管不确定性的举措都可能有利于 IPO。

9月14日周一
  1. Gary Marcus:The Road to AI We Can Trust(RSS)35

    Gary Marcus 评特朗普 9 月 24 日与中国谈 AI 的抉择

    Gary Marcus 在 The Economist 撰文提出,特朗普与习近平 9 月 24 日通话将把 AI 列入议程,他认为这可能是特朗普任内最具影响的决定,主张美中不应只谈芯片交易,而应就"AI 向善"寻求合作路径。他同时提到,当前 AI 股票下跌、公众反 AI 情绪升温,部分前盟友如 Steve Bannon 已转向反对阵营,若市场与民调继续走低,特朗普的立场可能生变。

  2. Gary Marcus:The Road to AI We Can Trust(RSS)67

    Gary Marcus 点评 Dario Amodei 的 AI 减速提案:三分肯定、七分质疑

    Gary Marcus 评 Dario Amodei 呼吁给 AI 发展减速的文章,Sam Altman 与 Elon Musk 已表态支持。Marcus 肯定其透明度承诺,但质疑其依赖与 AI 公司关系密切的 METR 做评估有监管捕获之嫌,指其拿中国当挡箭牌有损合作对话,并提出追责和产品召回等替代政策选项。文末提到特朗普反对减速,认为美国必须赢下 AI 竞赛。

    推荐理由:Gary Marcus 对 Dario Amodei 的减速提案给出有保留的支持,并指出监管捕获、追责与召回等被绕开的政策选项。

9月13日周日
  1. Peter McCrory68

    Peter McCrory 转发并推荐 Dario Amodei 的新文章《We Must Pace the Frontier》,该文主张 AI 行业应放慢速度并提出三步计划,Anthropic 单方面承诺其中第一步,即向第三方评估者提供永久、员工级别的系统访问权限,以便核查安全措施、报告事故和评估训练中的模型对齐。

    引用Dario Amodei@DarioAmodei

    We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier

  2. Aidan Gomez56

    Cohere CEO Aidan Gomez 引用 Sam Altman 关于认同 Dario 前沿限速、承诺接受独立评估机构的推文,并以讽刺口吻逐条批评:要求对手开放员工级访问、以安全为由关停不够安全的竞争者,以及中国不遵守就断供芯片。作者称这些想法是卡特尔式的做法。

    引用Sam Altman@sama

    I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.

  3. Jakub Pachocki71

    OpenAI 首席科学家 Jakub Pachocki 以一个爱心符号转发了 Dario Amodei 的新文章《We Must Pace the Frontier》,后者主张 AI 行业应放慢速度,并提出三部分计划,Anthropic 单方面承诺其中第一步,向第三方评估者提供永久、员工级别的系统访问权限,用于验证安全措施执行、报告事故并评估模型训练期间的对齐情况。全文见 https://darioamodei.com/post/we-must-pace-the-frontier。

    引用Dario Amodei@DarioAmodei

    We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier

9月12日周六
  1. Thinking Machines42

    我们自己的 @johnschulman2 与 Dwarkesh 对话,讨论随着模型不断进步和自我改进,人类判断力在哪些方面仍然重要:教它们处理混乱的现实世界任务,用品味判断什么在长期内有效,以及最重要的——明确我们真正想要什么。

    引用Dwarkesh Patel@dwarkesh_sp

    New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines

  2. Peter McCrory37

    这是该模型的一个重要局限。我们聚焦于 AI 转型的供给侧(AI 能做什么、扩散多快、工人转岗多快)。 价格是灵活的,总需求等于经济体的产出能力。 更多思考见 🧵

    引用modest proposal@modestproposal1

    Anthropic's economic scenario analysis is interesting. But this is not something you can ignore, this is the most important consideration! "the model cannot generate the negative feedback in which disruption depresses demand and amplifies its own labor-market consequences"

9月10日周四
  1. Peter McCrory52

    Anthropic 首席经济学家 Peter McCrory 与 Jack Clark 对谈其 AI 经济影响情景研究。他表示目标不是做预测,而是理解可能结果的区间及其出现的条件,希望厘清对不确定未来的分歧来源;引用内容提到研究情景从影响很小到 2030 年 GDP 增长 15%、知识工作者失业率达 18%。

    引用John Burn-Murdoch@jburnmurdoch

    New from us: Anthropic just published scenarios for AI’s possible economic impacts, which range from minimal, to explosive GDP growth of 15% by 2030 as knowledge-worker unemployment hits 18%. I sat down with their co-founder Jack Clark to pick his brains on how they’re thinking about all of this.

  2. jietang26

    你确定吗?找到最优模型规模很棘手:数据量、激活参数量、环境数量,以及目标推理成本。模型性能还取决于许多其他因素,每个因素都带来各自的变数。

    引用Charlie O'Neill@oneill_c

    Fable is probably ~2-2.5T parameters, not 10T. Kimi K3 is 2.8T params, trained on maybe 20–30k Blackwell-equivalents. It lands within spitting distance of Fable 5 in terms of capabilities (5, not 5.1). Anthropic has far more compute than Moonshot, better rl environments, better architecture and better optimizers and all of that adds to capability per parameter. So if Fable is only slightly ahead of K3 with this in mind, it's almost certainly a smaller model. GPT-5.5 and 5.6 are smaller still (I'll say more on that later)

  3. Peter McCrory47

    很好,与我们今天分享的内容互为补充。 评估决定 AI 在未来数年对增长影响大小的关键经济力量(并判断我们如今可能处于哪条路径上)是至关重要的工作。 干得漂亮 @alexolegimas @ben_moll

    引用Alex Imas@alexolegimas

    New post on the blog, featuring the excellent @ben_moll There’s been tons of discourse on how AI will contribute to economic growth, with many people closest to the technology predicting double digit increases. Are these forecasts likely? Probably not. The blog goes through the economics for why exploding improvements in capabilities (which technologists have been largely right about) may not translate to explosive growth. Ben’s thread covers this in detail, but gist is that: 1) there is nothing in economic growth models that prevents AI from leading to explosive growth but 2) this trajectory relies on a series of assumptions that are unlikely to hold in the real world. For example, one assumptions is likely to be violated because of a pretty counterintuitive feature of structural change: the sectors that become automated become smaller parts of the economy (because they’re cheaper, people become richer, and spending moves to non-automated parts of the economy). This, plus other features of the economy, is what will likely cause the trend of huge increases in capabilities coupled with “only” 4-5% growth (which is huge, btw) to continue. Here is the link: https://aleximas.substack.com/p/will-ai-soon-lead-to-double-digit Looking forward to hearing thoughts/feedback!

9月9日周三
  1. Together AI 研究与产品博客(RSS)67

    Together AI 深入解析开源 AI 编码栈:用 MIGHT 框架从闭源模型迁移到开源模型

    Together AI 发布长文,解析开发者从闭源模型转向开源模型所需的 AI 编码栈,提出由模型、推理、网关与路由、Harness、工具(Skills 与 MCP)组成的 MIGHT 五层框架。

    推荐理由:Together AI 把开源编码栈拆为 MIGHT 五层,给出大小模型分工和分层组合的具体实践方法。

  2. Mark Chen38

    两件事要区分: 在 Navier Stokes 工作中,有任何人类或智能体查看过用户数据吗?没有。 我们是否以整体方式使用用户反馈和去标识化数据来改进 ChatGPT 和 Codex?是的。每一家 LLM 公司都是如此。

    引用levent@__alpoge__

    “we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” i mean props to them for straight coming clean. (so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan) so i’ll now give a bit on my thinking here. i actually woulda been pumped to collaborate on this, there are a lot of people at oai i like (ok, clearly some were indirectly dicks to me because of being part of the whole situation, but im a big boy, i still like them), idgaf about authorship on that step anyway, coulda been me Tristan and every fte at oai for all i care (on that Tristan would disagree:p). but on hearing the loud convo in the hallway, especially the part where a millennium prize was offered if i’d just be removed from the paper, it was kinda clear the die had been cast and things were locked. pretty wacky, unstrategic, and unnecessary, since on my side things were mostly me and claude having a good time yoloing random stuff in the corner rather than anything institutional. i also like the idea of the labs cooperating, and even better on scientific progress. it’s a shame!

9月7日周一
  1. Mark Chen47

    同意 @JensenHuang 的观点:我们正在进入 AGI 时代。 AGI 时代也必须是对齐时代。我们需要教会 AI 热爱人类,并训练出与它们所监督的 AI 同样强大的 AI 监督者。 @merettm 在这篇深思熟虑、发人深省的文章中说得最好:https://openai.com/index/an-alien-mind

    引用Jensen Huang@JensenHuang

    @ChaseLochmiller @OpenAI GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years. AGI has arrived. Congratulations @OpenAI team. 400K GPUs coming online next.

9月6日周日
9月4日周五
9月3日周四
9月2日周三
  1. Microsoft:Official Blog(RSS)46

    微软:AI 基础设施的“良率”命题——从算力投入转向有用智能产出

    微软提出 AI 基础设施的“良率命题”,主张衡量标准应从建了多少算力转向产出多少有用智能。文中指出单个智能体任务消耗的 token 可达普通对话的 3400 倍以上,而全球 AI 渗透率仅为劳动人口的 18%,且以聊天为主。微软认为内存、网络与功耗的瓶颈需通过跨层协同设计解决,而非在单层堆叠资源。

  2. Jakub Pachocki53

    OpenAI 首席科学家 Jakub Pachocki 发文,试图纠正错误报道引发的对不可监控性的竞速担忧,指出包括 Astra 在内的当前前沿模型计算图深度与 GPT-4 相差不到两倍。他表示 OpenAI 从最早的推理模型起就保留并利用链式思维监控,认为该技术脆弱且趋势向坏,但与架构变化无关的原因将另行撰文说明,强化该技术是其当前研究项目的核心目标。

9月1日周二
  1. MIT News(RSS)25

    MIT 博士生 Ila Kumar:以社区共创方式设计 AI 与心理健康技术

    MIT 终身幼儿园小组博士生 Ila Kumar 主张社区共创式设计,让经历童年创伤、涉入儿童福利系统的年轻人从设计之初就参与技术开发。她与 Stepping Forward LA 合作开发以视觉拼贴替代文字沟通的应用,并与 Justice Resource Institute 合作设计支持青少年参与自身治疗计划制定的移动应用。

8月21日周五
  1. jietang46

    精彩评论:FLOPs 是智能;参数是知识!

    引用Liam Fedus@LiamFedus

    An excellent history of scaling laws from @jietang. In 2020, we explored the limits of sparsity in Switch Transformers by routing each token to only 1 out of 2048 experts (in retrospect, a bold choice). The model had fewer than 3B activated parameters, but 1.6T total parameters (comparable to today's frontier models). The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE. The lesson was that the optimal tokens-per-parameter ratio is highly task-dependent. Or as @NShazeer had already intuited: FLOPs were intelligence; parameters were knowledge!

8月20日周四
8月19日周三