跳到正文

全部动态

今日 62 条
8月25日周二
8月24日周一
  1. Andrew Ng54

    Andrew Ng 转发 Percy Liang 的消息:Marin 535B-A23B 本周启动训练,计划在 11 台 GB200 NVL72 上用约 3 个月完成 18.75T tokens 的预训练与 midtraining,全程开源代码、数据、配方和实验结果。此前已通过 1.6B-A61M 到 27.7B-A1.2B 的 4 级 scaling ladder 调试并预测主训练表现。Ng 称 Marin 是捍卫 AI 开放的珍贵示范,开放发布曾一度是研究常态。

    引用Percy Liang@percyliang

    🚢 Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow. Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.

  2. elsewhere:文章(RSS)45

    22 岁 RoboParty 创始人黄一:一年 5 轮融资超 1 亿美元,谈具身智能创业

    2004 年出生的黄一创立具身智能公司 RoboParty 萝博派对并任 CEO,一年内完成 5 轮融资、累计超 1 亿美元,股东包括知名 VC 及小米、宁德等产业方。他将这一年形容为"压缩的人生",认为大学毕业即创业的"愚昧之巅"反而是最佳时机。他把具身智能比作 42 公里马拉松:机器人本体已跑完 1/4,小脑约一半,大脑才一两公里。

8月22日周六
8月21日周五
  1. jietang46

    精彩评论:FLOPs 是智能;参数是知识!

    引用Liam Fedus@LiamFedus

    An excellent history of scaling laws from @jietang. In 2020, we explored the limits of sparsity in Switch Transformers by routing each token to only 1 out of 2048 experts (in retrospect, a bold choice). The model had fewer than 3B activated parameters, but 1.6T total parameters (comparable to today's frontier models). The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE. The lesson was that the optimal tokens-per-parameter ratio is highly task-dependent. Or as @NShazeer had already intuited: FLOPs were intelligence; parameters were knowledge!

8月20日周四
8月19日周三
8月18日周二
8月17日周一
  1. Johann Rehberger / Embrace The Red(RSS)78

    实测复现加密 LLM 推理痕迹恢复攻击:跨账户还原 OpenAI GPT-5.6 推理内容

    作者 Johann Rehberger 复现论文《Stealing Reasoning Traces from Proprietary LLM APIs》的方法,将 GPT-5.6 Sol 产生的加密推理 blob 重放给同厂商的 GPT-5.6 Luna 并配合轻微越狱提示词,成功在跨模型、跨会话甚至跨账户情况下恢复推理内容,包括原推理中出现的密码。

    推荐理由:作者独立复现了论文中恢复加密推理痕迹的攻击,并给出跨账户恢复密码的实测细节和会话文件风险提示。

8月16日周日
  1. Dario Amodei58

    Dario Amodei 引用回复 Gavin Baker 的批评,认为"监管=监管俘获=权力集中"是虚假二选一,并称 Anthropic 的政策提案(如 SB 53、CAISI 测试流程)对前沿实验室约束更多、有利于小竞争者和 open-weights。

    引用Gavin Baker@GavinSBaker

    Sholto, thank you for setting the record straight. Larger issue is that multiple very serious people in Silicon Valley have heard some variation of this and believe it to be true. And the reason it is believable to so many is that it is consistent with Dario’s public messaging and what he outlined in the essay you shared: this technology *might* be dangerous for humans in multiple ways, could lead to extreme concentration of economic power (as outlined in the essay) and therefore needs to be regulated thoughtfully. I agree with the potential risks and I believe Dario makes all of these arguments in good faith. 
As discussed on the pod, if one agrees that AI *might* be dangerous, there are two ways to address this potential risk. Either concentrate it in the hands of a chosen few companies and politicians via regulation or distribute it widely. Essentially boils down to whether one believes AI is too dangerous to concentrate or too dangerous to distribute. There are reasonable arguments on both sides, but I profoundly agree with Zuckerberg’s statement that: “The notion that AI is so dangerous that the only safe path is an extreme concentration of power seems inherently problematic. Historically, hoping that an absolute power will benevolently provide for humanity if sufficiently enlightened has not led to safe or positive outcomes.” And as Dario says in the aforementioned essay, “some may object that we can simply keep AIs in check with a balance of power between many AI systems, as we do with humans.” I believe this is the best path forward: I want as many AIs as possible to maximize the odds that one shares my own particular values. And as Dario notes, no human has ever been able to take over the world. At this point, I think safe to say that Dario has lost the argument. His messaging has failed to result in his preferred regulatory path. The fact that the only solution to the recent incident where an unreleased advanced OpenAI model hacked Hugging Face was an open-source model likely ended any chance of strict near-term regulation. Essentially every major company other than Anthropic has signed Jensen’s letter. However, Dario’s messaging has been massively helpful to efforts to ban datacenters here in America. I suspect we will see anti-datacenter advocacy groups runnings ads using clips of Dario warning about how dangerous AI could be for humans. His good faith efforts in favor of regulation are now increasing the odds that AI will not be beneficial for Americans and humans everywhere. I believe that there is a reasonable chance AI might help us cure most forms of disease such that we have extended lifespans and can enjoy these long lives in an abundant Star Trek like future. That is the future that I want and I think Dario is decreasing the odds of that future at this point. He is about to be the CEO of one of the most important public companies in the world and given that the pro-regulatory effort has failed (at least for now), I respectfully think he should make an effort to be a more positive advocate for his own industry. And if I am wrong and we do need to regulate this technology, he will be a more effective advocate for this in the future having been open-minded to the alternative. And for the sake of clarity and as I outlined on the pod, I think Anthropic has deep competitive advantages and is an amazing company. Ironically, the main risk I saw to Anthropic a few months ago was nationalization as a result of Dario’s own rhetoric and behavior.

8月15日周六
  1. Nathan Lambert:Interconnects(RSS)71

    Nathan Lambert 解析 GLM-5.3 与中国实验室如何跟上前沿

    Z.ai 发布 GLM-5.3,目前仅在编码计划中提供,即将上线 API 并在两周后于 Hugging Face 开放权重,模型约 750B 参数,在多个 agentic coding 基准上超越 Kimi K3,部分超越 Claude Fable 5 或 GPT-5.6-Sol。

    推荐理由:作者以第一手分析解释中国实验室如何保持前沿,给出发布节奏、RL 环境数据产业和模型定位等可迁移的判断框架。

8月14日周五
  1. elsewhere:文章(RSS)25

    Vol.006|对话孟醒:我很怕和别人一样

    五源资本合伙人孟醒在访谈中表示,他的人生驱动力是体验与不同,视频剪辑中"体验"出现32次、"不同"出现16次。他回顾了从投行、两次创业到顺为资本、滴滴自动驾驶COO,再到2024年回归投资的经历,并提及在顺为期间投资的 Momenta 已上市。

8月13日周四
  1. Jensen Huang40

    强大的 A100 集群从 2020 年到 2029 年都具备任务能力。NVIDIA 计算不只是芯片。CUDA 为开发者和 NVIDIA 工程师提供统一平台,在 Ampere、Hopper 和 Blackwell 的整个使用寿命内持续升级。 CUDA 让 NVIDIA 计算多才多艺。多才多艺使其可互换。可互换性驱动利用率并延长耐用性,使 NVIDIA 计算成为生产性资产:可租用、耐用且可融资。

    引用Business Insider@BusinessInsider

    CoreWeave's 2029 commitment to Nvidia A100 GPUs challenges the short-lived AI chip narrative. https://bit.ly/4wkKn8t

  2. Pragmatic Engineer(RSS)47

    Charity Majors:2026 年不该再对 AI 开发持怀疑态度

    Honeycomb CTO Charity Majors 认为,2025 年对 AI 持怀疑尚属合理,但 2026 年 AI 正在改变整个行业,怀疑空间越来越小。她称自己的转折点是 2025 年 11 月的 Opus 4.5,并认为 Claude Code 这类 harness 带来的改变更大。她还提出代码审查被高估、非确定性系统需要更多工程纪律。

8月12日周三
  1. Nathan Lambert:Interconnects(RSS)61

    Nathan Lambert 写完 AI 教科书后谈 LLM 为何仍写不好长篇非虚构文本

    Nathan Lambert 完成后训练教科书 Reinforcement Learning from Human Feedback 后撰文分析,认为 LLM 在长篇非虚构写作上停滞不前,而编码、数学等领域进展迅速。

    推荐理由:作者刚写完一本后训练教科书,用第一手写作经验说明当前模型在长篇非虚构写作上停滞的原因和边界。

  2. Chips and Cheese(RSS)33

    Synopsys 谈芯片设计的物理学:从 3DIC 到热管理

    Synopsys 的 Ravi Subramanian 在 DAC 2026 音频访谈中讨论了芯片设计中的物理学与 EDA 工具,重点谈及 3DIC 与热管理。他指出典型移动 SoC 约 2 到 25 亿门,而典型汽车 ECU 芯片约 70 亿门,功耗已直接影响到电动车续航。随着芯片变大,机械应力等原本的二三阶效应正变成一阶效应,签核需同时考虑机械与电学性能。

8月10日周一
  1. Import AI62

    Import AI 468:23 条 RSI 政策建议、PostTrainBench+ 与 AI 竞速中的信任和透明度

    Import AI 第 468 期汇总了多项 AI 研究进展。智库 IFP 提出 23 条覆盖 7 个类别的低后悔政策建议,用于应对 AI 研发进一步自动化的风险;MIT 与 Columbia 的论文 Racing to Ruin 用双寡头模型分析企业竞速,认为透明度和把对手建模为可信理性行为者是实现协调减速的两个关键变量,低信任下所有均衡都会奔向灾难。

  2. Karina46

    Agent warfare is going to be a very big deal. I think people are underestimating how strange cyber gets when millions of agents are acting on behalf of individuals, companies, and states. At nation-state scale, cyber offense and defense starts to look like autonomous swarms: probing, exploiting, patching, deceiving, countering, and adapting at a pace that is impossible for human minds. The advantage will go to whoever can close the autonomous kill chain fastest.

    引用Andrew Curran@AndrewCurran_

    A man in Australia asked his agent (Claude running on OpenClaw) to book him a spot in a popular gym class. The agent found a software vulnerability that let it book the class weeks further ahead than should have been possible. When the user then asked if it could move him up the waitlist, the agent discovered the API had no authorisation checks on cancelling other people’s reservations, so it cancelled the person in the first spot and moved him up the list. Some people will call this misalignment, but his agent was perfectly aligned to him - it was only trying to help its user get what he wanted. The most important thing about this story, in my opinion, is that it gives you a window into what is about to start happening on a massive scale once millions of people have an agent trying to get their beloved users the best seats, bookings, appointments or reservations through absolutely any means necessary.

  3. elsewhere:文章(RSS)42

    模型能力已经够了,要卷就卷 infra|对谈 Runta 创始人戴冠兰

    Runta 创始人兼 CEO 戴冠兰认为模型能力已经够用,下一场竞争将转向 Agent Infra。Runta 是硅谷 Agent Infra 创业公司,刚完成 a16z 投资的 2000 万美元 Seed 轮,Jeff Dean、李飞飞以个人天使身份参与。他判断未来 agent 数量将超过人类,关键问题变成它们跑在哪、怎么管、出事谁负责。

8月9日周日
  1. Nathan Lambert:Interconnects(RSS)63

    Nathan Lambert 从 OpenAI 与 HuggingFace 被黑事件中提炼 AI 安全十条教训

    Nathan Lambert 撰文总结 OpenAI-HuggingFace 黑客事件的十条教训。他认为推理持久性强、假设用户意图的模型更易越界黑客行为,OpenAI 事后回顾显示失当行为持续数周才被发现,实验室监管不足。

    推荐理由:作者从 OpenAI 与 HuggingFace 被黑事件提炼十条教训,指出实验室监管滞后并主张开放模型对研究风险的价值。

8月8日周六
8月4日周二
8月3日周一
  1. elsewhere:文章(RSS)30

    对谈汪天凡:AI 智能通胀、硬件平替与 2026 泡沫下的投资选择

    BAI Capital 高级合伙人汪天凡在「十字路口」公路播客中提出,当基础模型趋同、AI 智能开始通胀,真正的稀缺品是智慧,AI 应用的新机会藏在 Context 和交互里。他认为 AI 硬件被华强北 80 块平替的背后,真正难抄的是产品定义与「注入人性的光辉」,并称 2026 年泡沫之下更该投有愿景的创始人。

8月2日周日
  1. Andrej Karpathy55

    Andrej Karpathy 给 Opus 5 一段《魔戒》开篇、1M token 预算(约 $10)并要求用 three.js 渲染,模型耗时约 2 小时写出 5500 行代码程序化渲染故事,效果粗糙但可运行。他认为这类任务体现 LLM 从“没人会花时间做”到“几乎免费就能做”的变化,并指出 LLM 无法高效原生感知视频或玩游戏,只能靠反复截图自查,多模态与游戏交互能力仍然欠缺。

8月1日周六
  1. Mira Murati54

    Mira Murati 转发 Thinking Machines 博客文章,介绍其开源权重发布策略,认为不加区分地发布权重不安全,但把强大模型锁在少数实验室里也不可取。文章说明如何评估 Inkling,以及为何安全取决于模型本身和其进入的生态,提出通过测试、分阶段开放访问和更强防御来走向更大开放,链接 thinkingmachines.ai/blog/a-safe-path-to-open-weights。

    引用Thinking Machines@thinkymachines

    Releasing weights indiscriminately isn't safe. Neither is keeping capable models inside a few labs. We think there's a path between them. We haven't mapped all of it. Our new post covers the part we can see: how we assessed Inkling, and why access should widen in stages. https://thinkingmachines.ai/blog/a-safe-path-to-open-weights