New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines
#推理
#推理
今日 85 条
Thinking Machines@thinkymachinesAI 评分4242引用Dwarkesh Patel@dwarkesh_spDwarkesh Patel:Podcast & Blog(RSS)精选AI 评分6161 Dwarkesh 对谈 John Schulman、Beren Millidge 与 Charlie O'Neill:AI 研究者激辩递归自我改进还有多远
Dwarkesh Patel 邀请 Zyphra CTO Beren Millidge、Thinking Machines 首席科学家 John Schulman 和 Baseten 模型训练负责人 Charlie O'Neill 对谈递归自我改进(RSI)何时到来。
推荐理由:三位一线研究者围绕递归自我改进给出了各自不同的技术瓶颈判断,涵盖蒸馏、sim-to-real 与持续学习等具体分歧。
NVIDIA Technical Blog(开发者技术博客 · RSS)AI 评分2929 全栈 NIM 优化如何在 Nemotron 3 Ultra 上支撑 2.5 倍并发用户
NVIDIA 通过全栈 NIM 优化,在 Nemotron 3 Ultra 上实现 2.5 倍并发用户量。该优化针对生产环境部署大语言模型时,在现有 GPU 基础设施上提升并发服务能力并保持交互响应速度的需求,对提示词长、上下文跨步骤复用的智能体 AI 工作负载尤为关键。
SiliconFlow@SiliconFlowAI精选AI 评分7070推荐理由:上线方直接给出参数结构、上下文窗口、KV cache 对比和许可证信息,读者可据此评估实际部署选型。
vLLM 官方博客(RSS)AI 评分4040 追踪瓶颈:在 AMD Instinct MI355X 上优化 MiniMax M3
MiniMax M3 在 AMD Instinct MI355X 上的 vLLM 服务性能大幅提升:SemiAnalysis InferenceX 基准显示,并发 32 时 MXFP8 标准服务从 109.1 升至 342.4 output tokens/s/GPU(3.14×),中位 TTFT 从 1.46 秒降至 0.67 秒。
Ahead of AI(RSS)AI 评分7070 Sebastian Raschka 解析 GPT-6 Astra 与循环 Transformer 及隐藏推理链传闻
Sebastian Raschka 撰文点评 OpenAI 新发布的 GPT-6 Astra,认为它是其迄今用过最强的模型,在 3D 渲染、动画和计算机使用上提升尤其明显,并详解计算机使用训练流程(据报道 OpenAI 采购数万台 Mac mini/Mac Studio 作为 macOS 训练环境)。
OpenAI:官网动态(RSS · 排除企业/客户案例)精选AI 评分6161 OpenAI 发布 GPT-6 Astra:面向工作场景的新一代模型
OpenAI 发布 GPT-6 Astra,定位为面向商务的最强模型。该模型具备高级推理和计算机使用能力,写作与设计判断也更强。
推荐理由:官方发布了面向商务场景的新模型,点出推理、计算机使用和写作设计判断三项能力方向。
Sierra:Blog(RSS)AI 评分5353 Sierra 开源 hyper-𝜏-bench 基准,评测模型构建智能体的能力
Sierra 发布并开源 hyper-𝜏-bench(论文名 𝜏^𝜏-bench),一个衡量模型自主构建客服智能体能力的长程评测。最佳配置 Claude Opus 5(max reasoning)在 Claude Code 中独立通过 23.9% 的保留评测任务,与深度上下文的工程师协作时达 82.2%。
Noam Brown@polynoamialAI 评分6868引用Sebastien Bubeck@SebastienBubeckI would like to clarify a few things: 1) The screenshot is my reaching out to Levent to coordinate our releases. I hope it’s clear from the message that we came in with the best possible intentions. 2) I never ever asked for Levent to be removed from authorship of his own work (as indicated by my text). I was surprised to learn during the call with Tristan that they had only solved Euler and not Navier-Stokes; after learning this we brainstormed possible paths forward. One option we discussed was that Tristan could be the lead author on a rewrite of OpenAI’s Navier-Stokes proof. It is in that context that I said “it would be simpler if Levent was not an Anthropic employee” because I felt it would be inappropriate for an Anthropic employee to author OpenAI’s work. Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic. Another option I wanted to propose (but got cut short) is to offer access to our internal model so that they could try to finish their proof and bridge the gap between Euler and NS. Again I did not know how to navigate giving access to internal OpenAI IP to an Anthropic employee. 3) To reiterate it plainly: as my text clearly indicates, and as I said during our call, OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler. In the call I was immediately met with a litany of slander, including direct threats that if we were to announce Navier-Stokes he would immediately go to the press with a barrage of unfounded accusations. I refuted all these accusations but he replied “there is nothing you can do, I simply do not trust you”. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering, which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve. I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.) 4) Overall, on a personal level, it was incredibly difficult to have these conversations. Levent refused to attend any of the meetings despite my repeated asking. As Sholto Douglas said, there will need to be coordination between Anthropic and OpenAI in the future; I felt I was doing a proxy negotiation with Anthropic while the Anthropic employee refused to directly participate.
Noam Brown@polynoamialAI 评分7575OpenAI 宣布用一组智能体和比 GPT-6 Astra 更强的下一代模型给出纳维-斯托克斯千禧年大奖难题的解,Noam Brown 确认该结果耗资数百万美元。
引用OpenAI@OpenAIWe’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
NVIDIA Technical Blog(开发者技术博客 · RSS)AI 评分3434 NVIDIA Jetson 如何部署与优化前沿推理模型
NVIDIA 发布技术指南,介绍如何在 Jetson 边缘设备上部署和优化具备多步推理能力的模型。此前这类模型体积过大,无法在边缘硬件本地运行,开发者只能将推理请求路由至数据中心,带来网络依赖、成本上升与数据外泄风险。该指南称这一限制正在被打破。
OpenAI:GitHub 新仓库AI 评分4646 OpenAI 发布 openai/PrimeGaps186:素数间隙至多 186 的条件 Lean 形式化与数值证书
OpenAI 在 GitHub 新建仓库 openai/PrimeGaps186,给出素数间隙至多 186 的条件 Lean 形式化及数值证书。该仓库以 Lean 形式化配合数值证书的方式,对素数间隙不超过 186 这一结论进行条件性验证。
NVIDIA Technical Blog(开发者技术博客 · RSS)AI 评分3434 利用投机解码协同设计 AI 模型以加速 LLM 推理
NVIDIA 技术博客发布 AI 模型协同设计系列第三篇,探讨如何用投机解码在保持精度的同时加速 LLM 推理。文章给出五条指南,用于在帕累托前沿上选择草稿长度与草稿机制。
Google AI Developers@googleaidevsAI 评分3939
Z.ai@Zai_org精选AI 评分7575智谱(Z.ai)发布 GLM-5.3-Flash,为原生多模态模型,拥有 1M-token 上下文窗口,参数规模 320B-A18B,以 MIT 许可开源权重。
推荐理由:官方公布了定价定位、1M 上下文、MIT 许可和国产芯片适配等关键细节,可帮助读者评估这款轻量模型的实际可用性。
Hugging Face:Blog(RSS)精选AI 评分6666 IBM 发布 Granite 4.2 推理模型系列并详解构建过程
IBM Granite 团队发布 Granite 4.2 推理模型系列,包含 3B、8B、30B 三个 dense 版本,基于 Granite-4.1 基座(约 15T tokens 预训练,上下文扩展到 512K),经 SFT 和多阶段 GRPO 强化学习训练,8B 和 30B 额外经历 SWE、终端、搜索三类真实环境 agentic RL。
推荐理由:官方完整披露了从预训练、SFT 到多阶段 GRPO 强化学习的训练细节,读者可以据此了解推理模型的完整构建流程。
SemiAnalysis 长文 RSS(RSS)AI 评分5151 SemiAnalysis 发布 AgentX 1.0 开源百万上下文多轮智能体推理基准
SemiAnalysis 发布 AgentX 1.0,称其为全球首个 Apache 2.0 开源、100 万上下文的多轮智能体编码推理基准,投入超过 300 万美元构建数据集。
SemiAnalysis 长文 RSS(RSS)AI 评分5757 SemiAnalysis 分析开源模型追赶闭源前沿的周期规律:每代耗时减半
SemiAnalysis 将 LLM 史分为早期扩展、推理和智能体三个时代,按时代分别用当时基准测算开源与闭源模型的综合能力分。
Hugging Face:Blog(RSS)精选AI 评分6666 Liquid AI 发布 LFM2.5-DSpark 草稿模型,推理最高提速 3.2 倍
Liquid AI 为 LFM2.5-1.2B-Instruct、LFM2.5-2.6B 和 LFM2.5-8B-A1B 发布 DSpark 草稿模型检查点,采用推测解码,GPU 上吞吐最高提升 3.18 倍,端侧最高 2.87 倍,输出质量不变。
推荐理由:原文给出各模型在 GPU 和端侧的实测加速数据与 llama.cpp、SGLang 接入命令,读者可直接评估是否用于部署。
jietang@jietangAI 评分5353
Karina@karinanguyenAI 评分3737Grok 4.6 在 DiligenceBench 金融测试中以约 52–53% 位列第 2,与 Claude Opus 5 基本持平,Sonnet 5 以 46.2% 落后。
Import AIAI 评分3232 Import AI 469:DiG-bench 游戏基准、RSI 模拟器与扎克伯格的技术悲观主义
Import AI 469 介绍了新基准 DiG-bench(Discovery in Games),用 70 款规则与目标均隐藏的手工游戏测试 AI 自主发现环境规律的能力,目前仅公开 21 款。
Johann Rehberger / Embrace The Red(RSS)精选AI 评分7878 实测复现加密 LLM 推理痕迹恢复攻击:跨账户还原 OpenAI GPT-5.6 推理内容
作者 Johann Rehberger 复现论文《Stealing Reasoning Traces from Proprietary LLM APIs》的方法,将 GPT-5.6 Sol 产生的加密推理 blob 重放给同厂商的 GPT-5.6 Luna 并配合轻微越狱提示词,成功在跨模型、跨会话甚至跨账户情况下恢复推理内容,包括原推理中出现的密码。
推荐理由:作者独立复现了论文中恢复加密推理痕迹的攻击,并给出跨账户恢复密码的实测细节和会话文件风险提示。
vLLM 官方博客(RSS)AI 评分4343 vLLM 推出 DSpark 自适应验证:按置信度调度投机解码预算
vLLM 在 PR #47808 中引入 DSpark 置信度调度验证(enable_adaptive_verification),让引擎逐步决定验证多少 draft token,而非按部署固定 num_speculative_tokens。
Microsoft Research 博客(RSS)AI 评分3737 MindTopo 揭示多模态大模型的空间推理能力短板
微软研究院推出 MindTopo 基准,从连续性、分离、顺序、包围、绳结五类拓扑关系评估多模态大模型的推理与规划能力。测试显示,模型在静态图像识别上表现明显优于交互式规划任务,失败多发生在规划阶段而非感知阶段,且整体远低于人类水平。图像与视频生成仅在单帧关系可见时偶有帮助,跨多步动作时难以维持拓扑约束。
Michael Truell@mntruell精选AI 评分6161引用SpaceXAI@SpaceXAIIntroducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price.
推荐理由:转发并补充了 Grok 4.6 在难度任务和知识工作上更强、兼顾低成本低速度的定位,可作了解该版本能力方向的参考。
Google Research精选AI 评分6868 Google Research 提出知识画像框架:召回而非编码是 LLM 事实性瓶颈
Google Research 发布知识画像框架及 WikiProfile 基准(2,150 条 Wikipedia 事实,每条配 10 个任务),评估 13 个 LLM 的编码、召回与识别。
推荐理由:原文用知识画像框架区分编码与召回失败,并给出思维可回收约40-65%事实的关键数据,帮助读者理解事实性瓶颈所在。
Dwarkesh Patel:Podcast & Blog(RSS)AI 评分5454 Dwarkesh 对谈 Ryan Greenblatt:AI 能自动化 AI 研究后会发生什么
Dwarkesh Patel 与 Redwood Research 首席科学家 Ryan Greenblatt 辩论递归自我改进:Greenblatt 认为一旦 AI 能自动化 AI R&D。
Nathan Lambert:Interconnects(RSS)AI 评分5353 Nathan Lambert 的 RLHF 与 LLM 后训练教材由 Manning 出版并开售
Nathan Lambert 宣布其 post-training 教材《Reinforcement Learning from Human Feedback: Aligning and Post-training LLMs》由 Manning 出版,Manning 和 Amazon US 已开售,Amazon UK 十月发售。
Nathan Lambert:Interconnects(RSS)精选AI 评分6363 Nathan Lambert 从 OpenAI 与 HuggingFace 被黑事件中提炼 AI 安全十条教训
Nathan Lambert 撰文总结 OpenAI-HuggingFace 黑客事件的十条教训。他认为推理持久性强、假设用户意图的模型更易越界黑客行为,OpenAI 事后回顾显示失当行为持续数周才被发现,实验室监管不足。
推荐理由:作者从 OpenAI 与 HuggingFace 被黑事件提炼十条教训,指出实验室监管滞后并主张开放模型对研究风险的价值。
vLLM 官方博客(RSS)AI 评分4242 vLLM 在 Qwen3.5 上实现 25K 总 TPS/GPU
vLLM 社区完成 Qwen3.5 混合注意力架构的分离式服务优化,在 GB200 NVL72 系统上实现超过 25K 总 TPS/GPU。
OpenAI:GitHub 新仓库AI 评分4343 OpenAI 发布 openai/ten-proofs 仓库
OpenAI 在 GitHub 上线 openai/ten-proofs 仓库,收录数学与理论计算机科学中十个证明的 Lean 证书。
Meta Engineering Blog(RSS)AI 评分4747 Meta 广告排序的多阶段序列模型:从用户序列到 LLM 式缩放定律
Meta 为广告排序提出多阶段序列模型,将离线用户建模与在线排序解耦,并引入稠密 tokenization 与 target-aware attention,使序列模型呈现可预测的 LLM 式缩放定律。该平台已带来 Instagram 转化率 6%、Facebook 转化率 3%、Facebook 广告点击 3.5% 的累计提升,并成为 Meta 生成式广告推荐模型 GEM 的核心组件。
Dwarkesh Patel:Podcast & Blog(RSS)AI 评分3232 为什么更聪明的 AI 模型可能将算力价格推高 10 倍
Dwarkesh Patel 认为,更聪明的 AI 模型可能将算力价格推高 10 倍。该内容为其上周所写文章的视频录制版,原文可在其博客查看。视频由 Mercury 赞助,其内置 AI Command 可自动归类交易并同步至 QuickBooks。
美团 LongCat:HuggingFace 新模型AI 评分6464 美团 LongCat 发布 LongCat-Flash-Lite-Sparse 稀疏注意力模型,原生支持 1M token 上下文
美团 LongCat 发布 LongCat-Flash-Lite-Sparse,这是一个非思考型 MoE 模型,总参数 69B,每 token 激活约 3B,用 LongCat Sparse Attention(LSA)替换稠密 MLA,原生支持最长 1M token 上下文。
Chips and Cheese(RSS)AI 评分5252 Intel 向 RosaicLabs 提供 Atom RTL,神秘 32-tile AMX x86 实现身份成谜
据 Reuters 和 SemiAccurate 报道,Intel 向 2026 年 5 月注册、由 Rivos 老将创办的 RosaicLabs 提供 Atom CPU 的 RTL 代码,这不同于以往的架构授权。
Dwarkesh Patel:Podcast & Blog(RSS)精选AI 评分6060 Dwarkesh Patel 分析未来几年算力价格为何可能涨 10 倍以上
Dwarkesh Patel 撰文分析 AI 算力未来几年可能变得贵 10 倍以上的原因。他指出 Anthropic 收入同比约 10 倍增长而算力仅约 3 倍增长。
推荐理由:作者用 Anthropic 收入与算力增速的缺口推算算力价格走向,给出一条理解未来算力成本的经济分析思路。
BAIR:Berkeley AI Research Blog精选AI 评分6262 从 CUDA 到 MLX:K-Search 如何把数十年内核经验带到 Apple Silicon
IBM Research 基于 UC Berkeley Sky Lab 的 K-Search 框架扩展出 MLX 后端,并设计结构化 CUDA-to-MLX 翻译层,让进化式内核搜索把已有 CUDA 内核当作知识库适配到 Apple Silicon。
推荐理由:K-Search 把 CUDA 内核优化经验迁移到 Apple Silicon,读者可了解跨平台内核搜索的方法与实测数据。
vLLM 官方博客(RSS)AI 评分4343 vLLM 与 Speculators 开源支持 P-EAGLE、DFlash、DSpark 三种并行草稿算法
vLLM 与 Speculators 为 P-EAGLE、DFlash、DSpark 三种并行草稿(parallel drafting)算法提供完整开源支持,相关模型已发布在 RedHatAI HuggingFace Hub 的 Speculators Collection 中。