跳到正文

#推理

今日 85 条
9月25日周五
  1. Sakana AI57

    Sakana AI 宣布 Jürgen Schmidhuber 出任首席科学顾问,与现职兼任,将参与公司的 RSI Lab。正文称其自 1987 年学位论文起开创递归自我改进方向,1990 年提出世界模型思路,2018 年与 CEO David Ha 合著的 World Models 论文使其广为人知。他将于 10 月下旬来日,在东京举办公开活动。

    引用Sakana AI@SakanaAILabs

    Sakana AI welcomes Jürgen Schmidhuber as Chief Scientific Advisor. https://sakana.ai/schmidhuber/ Sakana AI is incredibly proud to announce that Jürgen Schmidhuber, universally recognized as the father of modern AI, is officially joining Sakana AI as Chief Scientific Advisor. For nearly four decades, Jürgen has explored how machines can learn to learn. His foundational work in the 1990s drove core advancements in deep learning and established early frameworks for world models. Crucially, his pioneering innovations in meta-learning opened the very path toward recursive self-improvement. These ideas have already shaped our own research, from the Darwin Gödel Machine to The AI Scientist. Now Jürgen will help guide our newly formed RSI Lab, whose objective is to trigger a compounding cycle of scientific discovery aimed at improving machine intelligence. We are assembling a critical mass of world-class experts in Tokyo to make this a reality. Welcome, @SchmidhuberAI !

9月24日周四
  1. NVIDIA Technical Blog(开发者技术博客 · RSS)39

    NVIDIA 推出 NV-Reason-CT Open 3D CT VLM,面向放射科医生链式推理

    NVIDIA 发布 NV-Reason-CT Open,这是一个面向 3D CT 体积影像的开放视觉语言模型,支持放射科医生的链式推理(Chain-of-Thought)。现有前沿通用模型在体积影像上表现不佳,多数开放医疗 AI 模型也缺乏相应能力,而 3D CT 是临床信息最丰富、数据最密集的模态之一,此前一直未被现代 VLM 充分覆盖。

  2. Dario Amodei66

    Anthropic 宣布由 Claude 主导发现噬菌体 DNA 中一个此前未知的酶系统,其基因旁有类似 CRISPR 的重复 DNA 阵列,已知少数同类系统均能切割、复制和粘贴 DNA,可能代表新的基因编辑机制,功能尚待研究。

    引用Anthropic@AnthropicAI

    Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR. We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use. Read more: https://www.anthropic.com/news/claude-discovers-novel-enzyme-system

    推荐理由:原文给出了发现的具体过程和人机协作验证方式,可帮助读者判断 AI 驱动生物学发现的可行路径。

9月23日周三
  1. elsewhere:文章(RSS)72

    大模型的斩杀线斩的是谁:从小米 MiMo-V2.6 看智能成本前沿的位移

    十字路口Crossing 发文分析大模型的「斩杀线」现象,即模型在智能和成本两个维度同时被超越后失去被选择的理由。9 月 22 日小米发布并开源 MiMo-V2.6 系列三款模型,智能水平接近前代两倍而价格不变,与同日发布的 Grok 4.7 智能指数持平但单项任务成本仅 0.13 美元,约为后者的 1/21 到 1/29。

  2. Apple Machine Learning Research(RSS)45

    Apple 提出 probe guidance:用扩散模型冻结内部状态引导流匹配,刷新扩散语言模型无条件生成 SOTA

    Apple 研究团队提出 probe guidance,利用已有扩散模型冻结的内部状态构建引导信号,无需在推理时增加额外前向计算,即可让弱模型与强模型保持相近动态。该方法在连续扩散语言模型的无条件生成上刷新 SOTA,应用于 1.7B 扩散语言模型时持续提升多项选择题基准表现。研究还发现,传统 autoguidence 中的弱模型必须来自训练的低熵区域。

  3. Simon Willison 博客84

    Anthropic 发布 Claude Opus 5.5,OpenAI 同日推出 GPT-6 Sol 和 GPT-6 Luna 掀起新一轮价格战

    Anthropic 于9月22日发布 Claude Opus 5.5,约一小时后 OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna。GPT-6 Luna 价格降至 $0.10/M 输入、$0.50/M 输出,为 GPT-5.6 Luna 的一半;GPT-6 Sol 同样减半至 $2/$10。

    推荐理由:作者用自己实测的价格表和 pelican 测试对比了三款新模型,还发现 Opus 5.5 max 会想满输出上限,可直接参考。

  4. Greg Brockman79

    OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna,基于 GPT-6 Astra 的技术进展,提供更快、更实惠的模型以支持大规模工作,覆盖专业工作、事实性、编码、computer use 和对齐等能力。两者还通过更高效的缓存和推理降价,API 价格比 GPT‑5.6 促销价低 50%。

    引用OpenAI@OpenAI

    Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.

    推荐理由:原文给出了两个新模型的能力来源与 API 降价幅度,读者可据此评估在大规模任务中替代 GPT‑5.6 的成本。

9月22日周二
  1. Latent Space(RSS)79

    Xiaomi MiMo-V2.6-Pro 1T-A42B 登顶开源权重模型,训练仅花费约 $3M

    Latent Space AINews 汇总 2026/9/19-9/21 AI 动态,核心是 Xiaomi 发布 MiMo-V2.6-Pro(1.02T 总参数/42B 激活,MIT 许可),以 Artificial Analysis Intelligence Index 46 分成为新的开源权重榜首,成本为 $0.435/M 输入、$0.87/M 输出 token。

    推荐理由:除发布信息外还汇总了 RL 成本与训练细节,读者可以看到开源权重模型追赶闭源的具体路径。

  2. StepFun63

    阶跃星辰发布 Step 5 Preview,在 Artificial Analysis Intelligence Index 得 44 分,作者称其将智能-成本 Pareto 前沿外推,每任务成本约 $0.71。引用的 Artificial Analysis 评测称其为 600B 总参数、27B 激活的 MoE 模型,定价 $1/$2.70 每百万输入/输出 token,上下文窗口 1M token,支持文本、图像和视频输入,当前闭源权重,计划 10 月 15 日开放权重。

    引用Artificial Analysis@ArtificialAnlys

    StepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index, matching Kimi K3 (max) at ~2.8x lower cost per task, but trails peers on agentic evaluations Step 5 Preview is @StepFun_ai's new flagship model, with 600B total and 27B active parameters, succeeding Step 3.7 Flash (released May 2026). It scores 44 on the Intelligence Index, level with Kimi K3 (max) and just behind GLM-5.3 (max, 45) and Qwen3.8 Max (45) Key takeaways: ➤ Step 5 Preview costs ~2.8x less per Intelligence Index task than models at the same score. It costs ~$0.72 per task, against ~$2.00 for Kimi K3 (max) at the same score of 44 and ~$2.01 for GLM-5.3 (max) at 45. This is driven by pricing: at $1/$2.70 per 1M input/output tokens, it is priced below both on input and output. MiMo-V2.6-Pro is the one model that scores higher (46) at a lower cost per task ($0.13) ➤ Frontier reasoning is the standout strength, and where the jump from Step 3.7 Flash is largest. Step 5 Preview scores 46% on Humanity's Last Exam, in line with Kimi K3 (max, 47%), and 21% on CritPt, between Kimi K3 (23%) and GLM-5.3 (max, 19%). Both are up sharply from Step 3.7 Flash: +25 points on HLE and +19 points on CritPt ➤ Higher AA-Omniscience accuracy than GLM-5.3 at fewer parameters, but with more hallucination. At 600B total parameters, Step 5 Preview reaches 42% accuracy on AA-Omniscience, our benchmark measuring factual recall and hallucination, ahead of GLM-5.3 (max, 34%, 753B) and behind Kimi K3 (max, 48%, 2.8T). It attempts more questions than GLM-5.3 (68% vs 55%) and hallucinates more often when it does (43% vs 30%), landing at 16 on the AA-Omniscience Index, between GLM-5.3 (14) and Kimi K3 (20) ➤ Agentic evaluations are where Step 5 Preview lags peers at a similar Intelligence Index score. It scores 1,566 Elo on GDPval-AA, our primary evaluation for agentic performance, behind Qwen3.8 Max (1,668) and GLM-5.3 (max, 1,646). The gap holds on Terminal-Bench 4.0 (33% vs 39% and 42%), AA-Briefcase (1,432 Elo vs 1,640 and 1,525) and AutomationBench-AA (51% vs 56% and 62%) Key model details: ➤ Model Size: 600B total parameters, 27B active MoE model ➤ Context window: 1M tokens ➤ Multimodality: Text, image and video input, text output ➤ Pricing: $1/$2.70 per 1M input/output tokens, with cached input at $0.05/M ➤ Availability: StepFun first-party API, with open weights release planned for October 15th ➤ Licensing: Closed weights currently, with weights release planned for October 15th

  3. Jeff Dean35

    感谢这场精彩的讨论,@dawnsongtweets!

    引用Dawn Song@dawnsongtweets

    I had the great honor and pleasure of sitting down with @JeffDean for his first public talk since leaving Google, where he spent an extraordinary 27 years. Few people have shaped modern computing and AI as profoundly - from MapReduce and Bigtable to TensorFlow, Mixture-of-Experts, TPUs, and Gemini. Our conversation covered some of the biggest questions shaping the future of AI: • How do you recognize a foundational idea before everyone else does? • How do you choose a research problem worth spending 5 years on? • What can coding teach us about building better reasoning models? • What might recursive self-improvement (RSI) actually look like? • What happens when the scientific discovery loop itself becomes increasingly automated? (and how is Jeff’s new startup going to contribute in this space?) • As AI becomes increasingly autonomous, how do we keep it safe and secure? • What should the next generation of researchers be working on? Here are some key insights and highlights for anyone building the future of AI. 🧵1/8

9月21日周一
  1. swyx23

    Jev 播客明天上线 在 apple /youtube 订阅 @latentspacepod 感谢 @allenpark 和 Ke 促成此事! 引用推文核心要点:在联合发明 ChatGPT 之后,我一直在问自己:为什么超人类对话模型没有带来 AGI? 过去 2 年我一直在隐身模式下构建一种新的模型训练方式(RLCD),以及一种新型前沿 AI 模型,今天正式发布:Jev • 快 20-200 倍 • 便宜 40-400 倍(输出 token 免费) • 为决策优化的前沿可组合智能 据我所知,这是通往 AI 驱动的经济革命的最短路径

    引用Diogo Almeida@CompleteSkeptic

    After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution

9月19日周六
9月18日周五
9月17日周四
  1. NVIDIA Technical Blog(开发者技术博客 · RSS)30

    TensorRT Edge-LLM 在 Jetson AGX Thor 上以 6.4 倍速度完成 MLPerf Edge Agentic Benchmark

    NVIDIA 的 TensorRT Edge-LLM 在 Jetson AGX Thor 上完成 MLPerf Edge Agentic Benchmark,速度提升 6.4 倍。该方案面向从云端数据中心迁移至车辆、机器人等边缘设备的 AI 智能体,针对其多步骤工具调用、结果评估与长对话推理带来的边缘推理新需求,实现快速 token 生成。

9月16日周三
  1. Google Research29

    Google Research 提出 Retrieve-for-Train:用 RL 编译扩散模型绕过推理瓶颈,加速复杂 AI 搜索

    Google Research 提出 Retrieve-for-Train 框架,通过离线强化学习发现奖励对齐的查询扇出并编译为监督信号,再蒸馏进一个 53.9M 参数的扩散检索器,实现推理时单次非自回归的查询扇出。该方法在 Gemma3-4B 和 Qwen3-4B 上微调扇出语言模型,用集合级属性奖励评估整组结果,无需人工标注,也无需推理时的 CoT 思考 token。

9月15日周二
  1. MiniMax (official)36

    H3 越来越快了。⚡️ @sgl_project + VDN-H3 现在让 MiniMax H3 在 8× B200 上突破 2 倍实时去噪——预热后端到端 9.0s 生成 14.4s 的 768p 视频,未测得质量下降。 开放模型通过开放生态持续进化。🚀

    引用SGLang@sgl_project

    SGLang-Diffusion with VDN-H3 now generates 14.4s of 768p video in just 9.0s 🚀 On 8× B200, 8 step denoising takes just 6.9s, reaching over 2× real time. The 9.0s figure covers the full generation request after warmup. No measured quality regression versus dense 50-step H3 across 103 test prompts. 🧵

9月14日周一
9月13日周日
9月12日周六