Sakana AI welcomes Jürgen Schmidhuber as Chief Scientific Advisor. https://sakana.ai/schmidhuber/ Sakana AI is incredibly proud to announce that Jürgen Schmidhuber, universally recognized as the father of modern AI, is officially joining Sakana AI as Chief Scientific Advisor. For nearly four decades, Jürgen has explored how machines can learn to learn. His foundational work in the 1990s drove core advancements in deep learning and established early frameworks for world models. Crucially, his pioneering innovations in meta-learning opened the very path toward recursive self-improvement. These ideas have already shaped our own research, from the Darwin Gödel Machine to The AI Scientist. Now Jürgen will help guide our newly formed RSI Lab, whose objective is to trigger a compounding cycle of scientific discovery aimed at improving machine intelligence. We are assembling a critical mass of world-class experts in Tokyo to make this a reality. Welcome, @SchmidhuberAI !
#推理
#推理
今日 85 条
Sakana AI@SakanaAILabsAI 评分5757引用Sakana AI@SakanaAILabs
Hugging Face:Blog(RSS)AI 评分4141 Liquid AI 发布 LFM2.5-VL-DSpark 草稿模型,为 LFM2.5-VL-3B 加速视觉语言推理
Liquid AI 发布实验性 DSpark 草稿模型 LFM2.5-VL-DSpark,为视觉语言模型 LFM2.5-VL-3B 加入投机解码路径,设备端解码最高提速 3.13x、H100 上 2.66x,端到端最高提升 2.62x 和 2.27x。
NVIDIA Technical Blog(开发者技术博客 · RSS)AI 评分3939 NVIDIA 推出 NV-Reason-CT Open 3D CT VLM,面向放射科医生链式推理
NVIDIA 发布 NV-Reason-CT Open,这是一个面向 3D CT 体积影像的开放视觉语言模型,支持放射科医生的链式推理(Chain-of-Thought)。现有前沿通用模型在体积影像上表现不佳,多数开放医疗 AI 模型也缺乏相应能力,而 3D CT 是临床信息最丰富、数据最密集的模态之一,此前一直未被现代 VLM 充分覆盖。
Dario Amodei@DarioAmodei精选AI 评分6666引用Anthropic@AnthropicAIClaude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR. We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use. Read more: https://www.anthropic.com/news/claude-discovers-novel-enzyme-system
推荐理由:原文给出了发现的具体过程和人机协作验证方式,可帮助读者判断 AI 驱动生物学发现的可行路径。
Fuli Luo@_LuoFuliAI 评分4545elsewhere:文章(RSS)AI 评分7272 大模型的斩杀线斩的是谁:从小米 MiMo-V2.6 看智能成本前沿的位移
十字路口Crossing 发文分析大模型的「斩杀线」现象,即模型在智能和成本两个维度同时被超越后失去被选择的理由。9 月 22 日小米发布并开源 MiMo-V2.6 系列三款模型,智能水平接近前代两倍而价格不变,与同日发布的 Grok 4.7 智能指数持平但单项任务成本仅 0.13 美元,约为后者的 1/21 到 1/29。
Apple Machine Learning Research(RSS)AI 评分4545 Apple 提出 probe guidance:用扩散模型冻结内部状态引导流匹配,刷新扩散语言模型无条件生成 SOTA
Apple 研究团队提出 probe guidance,利用已有扩散模型冻结的内部状态构建引导信号,无需在推理时增加额外前向计算,即可让弱模型与强模型保持相近动态。该方法在连续扩散语言模型的无条件生成上刷新 SOTA,应用于 1.7B 扩散语言模型时持续提升多项选择题基准表现。研究还发现,传统 autoguidence 中的弱模型必须来自训练的低熵区域。
Simon Willison 博客精选AI 评分8484 Anthropic 发布 Claude Opus 5.5,OpenAI 同日推出 GPT-6 Sol 和 GPT-6 Luna 掀起新一轮价格战
Anthropic 于9月22日发布 Claude Opus 5.5,约一小时后 OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna。GPT-6 Luna 价格降至 $0.10/M 输入、$0.50/M 输出,为 GPT-5.6 Luna 的一半;GPT-6 Sol 同样减半至 $2/$10。
推荐理由:作者用自己实测的价格表和 pelican 测试对比了三款新模型,还发现 Opus 5.5 max 会想满输出上限,可直接参考。
Latent Space(RSS)AI 评分5252 Latent Space 访谈 John Platt:Google ERA 用 AI 求解可评分科学问题与气候应用
Latent Space 发布对 John Platt 的访谈,介绍其团队在 Google 的 Empirical Research Assistance(ERA)项目。
Greg Brockman@gdb精选AI 评分7979引用OpenAI@OpenAIPlease welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
推荐理由:原文给出了两个新模型的能力来源与 API 降价幅度,读者可据此评估在大规模任务中替代 GPT‑5.6 的成本。
Latent Space(RSS)精选AI 评分7979 Xiaomi MiMo-V2.6-Pro 1T-A42B 登顶开源权重模型,训练仅花费约 $3M
Latent Space AINews 汇总 2026/9/19-9/21 AI 动态,核心是 Xiaomi 发布 MiMo-V2.6-Pro(1.02T 总参数/42B 激活,MIT 许可),以 Artificial Analysis Intelligence Index 46 分成为新的开源权重榜首,成本为 $0.435/M 输入、$0.87/M 输出 token。
推荐理由:除发布信息外还汇总了 RL 成本与训练细节,读者可以看到开源权重模型追赶闭源的具体路径。
StepFun@StepFun_aiAI 评分6363引用Artificial Analysis@ArtificialAnlysStepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index, matching Kimi K3 (max) at ~2.8x lower cost per task, but trails peers on agentic evaluations Step 5 Preview is @StepFun_ai's new flagship model, with 600B total and 27B active parameters, succeeding Step 3.7 Flash (released May 2026). It scores 44 on the Intelligence Index, level with Kimi K3 (max) and just behind GLM-5.3 (max, 45) and Qwen3.8 Max (45) Key takeaways: ➤ Step 5 Preview costs ~2.8x less per Intelligence Index task than models at the same score. It costs ~$0.72 per task, against ~$2.00 for Kimi K3 (max) at the same score of 44 and ~$2.01 for GLM-5.3 (max) at 45. This is driven by pricing: at $1/$2.70 per 1M input/output tokens, it is priced below both on input and output. MiMo-V2.6-Pro is the one model that scores higher (46) at a lower cost per task ($0.13) ➤ Frontier reasoning is the standout strength, and where the jump from Step 3.7 Flash is largest. Step 5 Preview scores 46% on Humanity's Last Exam, in line with Kimi K3 (max, 47%), and 21% on CritPt, between Kimi K3 (23%) and GLM-5.3 (max, 19%). Both are up sharply from Step 3.7 Flash: +25 points on HLE and +19 points on CritPt ➤ Higher AA-Omniscience accuracy than GLM-5.3 at fewer parameters, but with more hallucination. At 600B total parameters, Step 5 Preview reaches 42% accuracy on AA-Omniscience, our benchmark measuring factual recall and hallucination, ahead of GLM-5.3 (max, 34%, 753B) and behind Kimi K3 (max, 48%, 2.8T). It attempts more questions than GLM-5.3 (68% vs 55%) and hallucinates more often when it does (43% vs 30%), landing at 16 on the AA-Omniscience Index, between GLM-5.3 (14) and Kimi K3 (20) ➤ Agentic evaluations are where Step 5 Preview lags peers at a similar Intelligence Index score. It scores 1,566 Elo on GDPval-AA, our primary evaluation for agentic performance, behind Qwen3.8 Max (1,668) and GLM-5.3 (max, 1,646). The gap holds on Terminal-Bench 4.0 (33% vs 39% and 42%), AA-Briefcase (1,432 Elo vs 1,640 and 1,525) and AutomationBench-AA (51% vs 56% and 62%) Key model details: ➤ Model Size: 600B total parameters, 27B active MoE model ➤ Context window: 1M tokens ➤ Multimodality: Text, image and video input, text output ➤ Pricing: $1/$2.70 per 1M input/output tokens, with cached input at $0.05/M ➤ Availability: StepFun first-party API, with open weights release planned for October 15th ➤ Licensing: Closed weights currently, with weights release planned for October 15th
Simon Willison 博客AI 评分7373 TypeSafe AI 发布决策模型 Jev,只返回带置信度的数值输出
TypeSafe AI 上周发布 Jev,称为 System One 模型,作者更倾向叫决策模型。它接受文本输入但输出浮点数形式的分类、是非判断、评分及置信度,只按输入收费,价格为 $0.042 per million tokens,低于 OpenAI GPT-5 Nano 的 $0.05/million,输出免费。
Fuli Luo@_LuoFuliAI 评分6666小米 MiMo 团队发布 MiMo-V2.6,称其可能是开源模型团队迄今计算量最大的单次 RL 运行之一,通过 mid-training 与高强度 RL 打造,现为排名第一的开源模型。
Latent Space(RSS)AI 评分6868 Latent Space 访谈 TypeSafe AI CEO Diogo Almeida:谈 Jev 与 System 1 模型
Latent Space 播客访谈 TypeSafe AI CEO Diogo Almeida,介绍其新发布的 Jev 模型,定位为面向软件而非聊天的 System 1 可编程模型,优化智能与成本之比。
Jeff Dean@JeffDeanAI 评分3535引用Dawn Song@dawnsongtweetsI had the great honor and pleasure of sitting down with @JeffDean for his first public talk since leaving Google, where he spent an extraordinary 27 years. Few people have shaped modern computing and AI as profoundly - from MapReduce and Bigtable to TensorFlow, Mixture-of-Experts, TPUs, and Gemini. Our conversation covered some of the biggest questions shaping the future of AI: • How do you recognize a foundational idea before everyone else does? • How do you choose a research problem worth spending 5 years on? • What can coding teach us about building better reasoning models? • What might recursive self-improvement (RSI) actually look like? • What happens when the scientific discovery loop itself becomes increasingly automated? (and how is Jeff’s new startup going to contribute in this space?) • As AI becomes increasingly autonomous, how do we keep it safe and secure? • What should the next generation of researchers be working on? Here are some key insights and highlights for anyone building the future of AI. 🧵1/8
eric zakariasson@ericzakariassonAI 评分7575
elsewhere:文章(RSS)AI 评分5050 清华叉院徐梦迪谈具身智能:从 In-Context Learning 到 Scaling Law 的信号
清华交叉信息研究院助理教授徐梦迪在播客访谈中提出,机器人泛化的核心路径是 In-Context Learning——通过一两次交互当场学会新任务,而非仅依赖预训练。
swyx@swyxAI 评分2323引用Diogo Almeida@CompleteSkepticAfter co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
vLLM 官方博客(RSS)AI 评分3535 vLLM 在 GB300 NVL72 上 PD 部署 Qwen3.8-2.4T 的性能结果
vLLM 公布 Qwen3.8-2.4T 在 GB300 NVL72 集群上的 PD 分离部署结果:8K/1K 负载下高吞吐场景达每 GPU 5000 total token 吞吐,低延迟场景每用户 180 生成 token,并给出完整 pareto 前沿。
Ethan Mollick:One Useful Thing(RSS)AI 评分5757 Ethan Mollick 谈 AI 能力过悬:GPT-6 Astra 与 Fable 5.1 已能完成数周人类工作,人类需靠四大优势与之协作
Ethan Mollick 撰文指出 AI 发展仍在指数曲线上,而人类系统跟上得很慢,现有模型能力与实际使用之间存在巨大"能力过悬"。
Epoch AI@EpochAIResearchAI 评分4949
SemiAnalysis 长文 RSS(RSS)AI 评分6161 SemiAnalysis 实测 DeepSeek-V4.1-Flash Engram 的 DRAM/SSD 卸载:HBM 带宽比容量更关键
SemiAnalysis 研究了 DeepSeek-V4.1-Flash 的 Engram 多 token 查找机制,其表可卸载到主机 DRAM,同质量下降低 HBM 需求。
Dwarkesh Patel:Podcast & Blog(RSS)精选AI 评分6363 Dwarkesh 对谈 Noam Brown:多智能体、对齐与递归自我改进
Dwarkesh Patel 播出与 OpenAI 研究员 Noam Brown 的对谈,Brown 是 o1 及推理模型的基础贡献者之一,现负责多智能体系统。
推荐理由:Noam Brown 亲述万级智能体协作的实测细节与对齐担忧,对理解推理模型下一步走向有直接参考价值。
SenseTime@SenseTime_AIAI 评分5353商汤发布 SenseNova U1.5 技术报告,这是一个开源的 8B-MoT 原生统一模型,通过共享注意力连接理解与生成。
NVIDIA Technical Blog(开发者技术博客 · RSS)AI 评分3030 TensorRT Edge-LLM 在 Jetson AGX Thor 上以 6.4 倍速度完成 MLPerf Edge Agentic Benchmark
NVIDIA 的 TensorRT Edge-LLM 在 Jetson AGX Thor 上完成 MLPerf Edge Agentic Benchmark,速度提升 6.4 倍。该方案面向从云端数据中心迁移至车辆、机器人等边缘设备的 AI 智能体,针对其多步骤工具调用、结果评估与长对话推理带来的边缘推理新需求,实现快速 token 生成。
Fuli Luo@_LuoFuliAI 评分3838
NVIDIA Blog(RSS)AI 评分5454 NVIDIA Vera Rubin NVL72 首次参加 MLPerf Inference v6.1,吞吐最高达 GB300 NVL72 的 3.7x
NVIDIA Vera Rubin NVL72 首次提交 MLPerf Inference v6.1 预览结果,在 Qwen3-VL 上吞吐最高达 GB300 NVL72 的 3.7x,在 DeepSeek-R1 上达 2.5x,分别使用 vLLM 搭配 NVIDIA Dynamo 和 TensorRT-LLM。
Apple Machine Learning Research(RSS)AI 评分4040 DACA-GRPO:面向扩散语言模型强化学习的去噪感知信用分配方法
Apple 研究者提出 DACA-GRPO,一种可插拔的 GRPO 训练增强方法,用于扩散语言模型强化学习。它通过 Denoising Progress Scores 提取逐 token 重要性权重(无额外前向开销),并用 Stratified Masking Likelihood 降低 mean-field 似然偏差。
Latent Space(RSS)AI 评分5050 Latent Space 访谈 Good Start Labs CEO:游戏训练能否迁移到真实工作
Latent Space 访谈 Good Start Labs CEO Alex Duffy,探讨用 Diplomacy、1830 等游戏训练 AI 模型能否让技能迁移到真实工作。
Google ResearchAI 评分2929 Google Research 提出 Retrieve-for-Train:用 RL 编译扩散模型绕过推理瓶颈,加速复杂 AI 搜索
Google Research 提出 Retrieve-for-Train 框架,通过离线强化学习发现奖励对齐的查询扇出并编译为监督信号,再蒸馏进一个 53.9M 参数的扩散检索器,实现推理时单次非自回归的查询扇出。该方法在 Gemma3-4B 和 Qwen3-4B 上微调扇出语言模型,用集合级属性奖励评估整组结果,无需人工标注,也无需推理时的 CoT 思考 token。
vLLM 官方博客(RSS)AI 评分4545 如何用 GB300 NVL72 为 Kimi-K3 训练最快的 DSpark 推测解码器
vLLM 团队借助 Speculators 训练库,为 2.8T 参数的 Kimi K3 训练出 DSpark 推测解码器,在数学推理上把单流交互速度从约 110 提升到约 435 tok/s/user,并发负载下同等交互性时输出吞吐最高提升约 3.5 倍。
MiniMax (official)@MiniMax_AIAI 评分3636引用SGLang@sgl_projectSGLang-Diffusion with VDN-H3 now generates 14.4s of 768p video in just 9.0s 🚀 On 8× B200, 8 step denoising takes just 6.9s, reaching over 2× real time. The 9.0s figure covers the full generation request after warmup. No measured quality regression versus dense 50-step H3 across 103 test prompts. 🧵
SemiAnalysis 长文 RSS(RSS)AI 评分5757 SemiAnalysis 实测 Vera Rubin NVL72 智能体推理:每美元 TCO 吞吐最高约 67 倍于 GB300
SemiAnalysis 在自建 AgentX 智能体推理基准上发布 Rubin 平台首批经核验的实测结果,称在 170 TPS、自有 TCO 口径下,Vera Rubin NVL72 相比 GB300 Dynamo TRTLLM 实现约 67 倍的总吞吐;在更常见的 60-100 TPS 区间为 1.4-3 倍。
蚂蚁 inclusionAI:HuggingFace 新模型AI 评分4848 蚂蚁 inclusionAI 发布 Qwen3.8-27B-singprobe 流式安全探针
蚂蚁 inclusionAI 发布 Qwen3.8-27B-singprobe,这是基于 Qwen/Qwen3.8-27B 的流式护栏探针,复用基座模型隐藏状态,逐 token 对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
蚂蚁 inclusionAI:HuggingFace 新模型AI 评分4747 蚂蚁 inclusionAI 发布 Qwen3.6-27B-singprobe 流式安全探针
蚂蚁 inclusionAI 开源 Qwen3.6-27B-singprobe,一个基于 Qwen/Qwen3.6-27B 的流式安全探针,复用基座模型隐藏状态,在每个 token 上对查询意图、回复不安全和幻觉风险打分,解码开销低于 0.5%。
蚂蚁 inclusionAI:HuggingFace 新模型AI 评分4949 蚂蚁 inclusionAI 发布 Qwen3.5-35B-A3B-singprobe 流式安全探针
蚂蚁 inclusionAI 在 HuggingFace 发布 Qwen3.5-35B-A3B-singprobe,一个基于 Qwen/Qwen3.5-35B-A3B 的流式安全探针,仅 4.2M 参数,复用基座模型隐藏状态逐 token 输出 8 类意图、不安全与幻觉风险评分,解码开销低于 0.5%。
SemiAnalysis 长文 RSS(RSS)AI 评分5454 SemiAnalysis 分析为何 4-hi HBM 是推理的最优配置
SemiAnalysis 长文论证下一代加速器正从 12-hi 转向 8-hi 乃至 4-hi HBM 堆叠,Nvidia Rubin Ultra 将单 GPU HBM 从 288GB 降至 192GB。
vLLM 官方博客(RSS)AI 评分5050 vLLM 中 Kimi K3 性能优化:吞吐量提升至 2.8 倍
vLLM 通过自适应投机 token 预算、KDA 前缀检查点、零拷贝混合 KDA 批次和延迟 MXFP4 收尾等优化,将 Kimi K3 服务性能从 v0.27.1 提升至 main:延迟降低 56%–60%,吞吐量提升 2.2–2.8 倍,TTFT 降低 72%–85%(并发 1、4、16,8K/1K 负载,TP8)。
Mira Murati@miramuratiAI 评分4242引用Thinking Machines@thinkymachinesOur own @johnschulman2 talks with Dwarkesh about where human judgment still matters as models improve and self-improve: teaching them to handle messy real-world tasks, applying taste to what works in the long run, and, above all, specifying what we actually want.