跳到正文

#推理

今日 85 条
7月27日周一
  1. vLLM 官方博客(RSS)72

    vLLM 发布 Kimi K3 Day-0 支持,2.8T MoE 模型最高 370 tok/s

    vLLM 宣布对月之暗面 Kimi K3 提供 Day-0 高效支持。Kimi K3 是 2.8 万亿参数的 MoE 多模态模型,每 token 激活 896 个专家中的 16 个,基于 Kimi Delta Attention 与 Attention Residuals,支持 1M token 上下文窗口。

    推荐理由:vLLM 官方详述 Kimi K3 的服务适配与内核优化,读者可以据此了解 2.8T 混合 MoE 的生产部署配方和性能数据。

7月26日周日
  1. BAIR:Berkeley AI Research Blog32

    Berkeley AI Research 提出 ABBEL:用信念状态监督让 LLM 高效长程交互

    Berkeley AI Research 提出 ABBEL 框架,将 LLM 智能体的摘要以自然语言"信念状态"形式隔离并监督其信息内容,替代完整交互历史作为工作上下文。在协作编程基准 CollabBench 上,采用重建式信念评分的 ABBEL-rec-BG 将自摘要模型与全上下文模型的性能差距缩小约 50%,训练步数从 100 降至 50,峰值上下文 token 长度也低于全上下文设置。

7月23日周四
7月20日周一
7月18日周六
7月17日周五
  1. Marc Andreessen 🇺🇸55

    Edgar Dobriban 在 AI(GPT-5.6 Sol Pro)协助下否定了二十年的猜想:Benjamini-Hochberg 方法在相关双侧高斯检验下并不总能将 false discovery rate 控制在名义水平,构造的因子模型在 alpha=0.01 下证明了 FDR>0.0104。GPT-5.6 用 90 分钟推理一次性解决问题,而 GPT-5.5 迭代约 20 小时未能解决;偏离幅度较小,意义主要是概念性的。预印本见 https://faculty.wharton.upenn.edu/wp-content/uploads/2017/06/bh.pdf,代码见 https://github.com/dobriban/BH。

    引用Edgar Dobriban@EdgarDobriban

    AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the Benjamini-Hochberg or BH method) and proved it controls the FDR. This method has been widely adopted in modern high-throughput science, including in genomics, astronomy, economics, etc. The paper has has garnered more than 130,000 citations to date. However Benjamini and Hochberg showed FDR control only when the data for the individual tests are *independent*. In practice, these data are often dependent; a good example is data on genetic variants due to linkage disequilibrium. Later work has focused on extending the validity of the BH procedure, e.g., to a form of positive dependence by Benjamini and Yekutieli (2001). The question of when the BH procedure controls the FDR has remained open. Over the last twenty years, many authors, including Reiner-Benaim (2007), Kim and van de Wiel (2008), Benjamini (2010), Sarkar (2023), Sarkar and Zhang (2025), have conjectured that the BH procedure controls the FDR for two-sided tests using any correlated Gaussian data. These authors have presented both theoretical and empirical evidence supporting, but not directly showing, the conjecture. With the help of AI (specifically GPT-5.6 Sol Pro), I have settled the question in the negative: The Benjamini-Hochberg procedure does *not* generally control the false discovery rate at the desired level for correlated two-sided Gaussian tests. This was done by exhibiting a Gaussian factor model for which, at a nominal level alpha=0.01, the false discovery rate is proved to be FDR>0.0104. There is a lot of interesting commentary to be made: 1. This result should be of interest to everybody in the field of statistics. Emmanuel Candes of Stanford University once called the false discovery rate and the Benjamini-Hochberg procedure "one of the two most important developments in statistics after 1950" (the other being James-Stein shrinkage). The present conjecture is probably the most central question about FDR/BH that was unresolved to date. 2. GPT-5.6 one-shot the problem after 90 minutes of reasoning, whereas with 5.5 I was not able to solve it even after iterating with multiple parallel agents for perhaps 20 hours. So the capability improvement is quite real. Exciting times to live in! 3. The argument is not especially surprising, but it does combine an asymptotic approach (standard for FDR analysis, see e.g., Genovese and Wasserman, Efron, etc) with a numerical certificate in a way that would be pretty non-standard in the field. Once we have the specific example, then straightforward simulations also support that the false discovery rate is indeed higher than the nominal value (see attached fig). 4. The current degree of violation over the nominal level is relatively small (0.104 vs 0.1). So the importance of this result is mainly conceptual. The practical implications remain to be determined. Overall, an exciting development! Preprint is available here (https://faculty.wharton.upenn.edu/wp-content/uploads/2017/06/bh.pdf) and will be on arxiv tonight; supporting code is here (https://github.com/dobriban/BH).

7月15日周三
  1. Hugging Face:Blog(RSS)81

    Thinking Machines 发布 1T 参数多模态开源模型 Inkling 及 Inkling-Small

    Thinking Machines 在 Hugging Face 上发布开源多模态模型 Inkling,共 975B 总参数、41B 激活参数,支持 1M 上下文,原生接收图像、文本和音频输入,训练数据为 45 万亿 token,并同时发布 276B 总参数、12B 激活参数的 Inkling-Small。

    推荐理由:文章给出了 Inkling 的架构细节、各档位 VRAM 需求和多条部署路径,便于读者评估在自己的硬件上运行哪种变体。

7月14日周二
7月13日周一
7月11日周六
  1. Dwarkesh Patel:Podcast & Blog(RSS)45

    Adam Brown 深入浅出讲解广义相对论:从爱因斯坦的“最快乐思想”到黑洞

    Google DeepMind BlueShift 负责人 Adam Brown 在播客中深入浅出讲解广义相对论,从爱因斯坦“最快乐的思想”出发,解释引力为何是弯曲时空的产物而非一种力,并讲解黑洞为何无法被用来提取无限能量。节目最后还讨论了 AI 距离从零重新发现广义相对论还有多远。

7月8日周三
  1. Hugging Face:Blog(RSS)63

    Hugging Face 发布 native-speed vLLM transformers 后端

    Hugging Face 宣布 vLLM 的 transformers 建模后端现在达到甚至超过 vLLM 原生实现的吞吐速度,模型作者无需移植代码即可用 --model-impl transformers 获得 vLLM 级推理性能。

    推荐理由:官方给出了与 vLLM 原生实现对比的具体吞吐数字和使用方法,读者可以据此判断是否切换到自己已有的 transformers 模型工作流。

7月7日周二
7月6日周一
7月4日周六
7月3日周五
7月2日周四
6月30日周二
  1. Dwarkesh Patel:Podcast & Blog(RSS)51

    Dwarkesh Patel 对谈 Grant Sanderson:AI 与数学的未来

    Dwarkesh Patel 发布与 3Blue1Brown 主理人 Grant Sanderson 的播客对谈,讨论 AI 在数学领域的快速进展及其对其他领域的预示。两人探讨了 IMO 表现、黎曼假设的可能证明形态、概念性突破的百年验证周期、AI 连接不同数学领域的潜力,以及现实任务为何难以纳入 RL 环境,并谈到数学家未来可能转向策展与解释角色。

6月29日周一
6月26日周五
  1. Dwarkesh Patel:Podcast & Blog(RSS)67

    Dwarkesh Patel:下一个突破将是 AI 在工作中学习

    Dwarkesh Patel 撰文提出,实验室押注 RLVR 在数千环境中训练可通向 AGI,但仅在会话内学习不够,真正的下一个范式是让 AI 从部署后的真实经验中持续学习并回写到权重。

    推荐理由:作者用计算机使用进度慢的例子说明可验证性之外还需可反复模拟性,并给出把部署经验蒸馏回权重的可能路径。

6月25日周四
6月12日周五
6月11日周四
  1. Google DeepMind:Blog(RSS)76

    Google DeepMind 发布开源实验模型 DiffusionGemma,GPU 文本生成最高提速 4 倍

    Google DeepMind 发布实验性开源模型 DiffusionGemma,基于文本扩散方法,在专用 GPU 上实现最高 4 倍的文本生成提速。该模型为 26B 总参数的 MoE,推理时仅激活 3.8B 参数,量化后可装入 18GB 显存,单张 NVIDIA H100 上超过 1000 tokens 每秒,RTX 5090 上超过 700 tokens 每秒。

    推荐理由:官方给出具体吞吐数字、显存占用和质量取舍,读者可据此判断它适不适合本地交互式工作流。

6月10日周三
  1. vLLM 官方博客(RSS)75

    vLLM 原生支持 Google DeepMind 26B 扩散语言模型 DiffusionGemma

    Google DeepMind 与 vLLM 团队合作,将 26B 参数、基于 Gemma4 的离散扩散语言模型 DiffusionGemma 接入 vLLM,这是 vLLM 首个原生支持的 dLLM。

    推荐理由:原文详解了扩散语言模型的解码机制和 vLLM 的 ModelState 接入方案,并给出 H200 上的吞吐数据,对推理工程实践有直接参考价值。

  2. Andrej Karpathy76

    Anthropic 发布 Claude Fable 5,作者称其与 Mythos 为同一底层模型但增加了安全护栏,基准几乎全面 SOTA。他认为质变程度堪比去年 11 月的 Claude 4.5,尤其擅长长程难题求解,但护栏目前过于敏感;软件唾手可得令他的软件需求大增(Jevon 悖论)。

    引用Claude@claudeai

    Fable 5 is state-of-the-art on nearly all tested benchmarks, with exceptional performance in software engineering, knowledge work, scientific research, and vision. The longer and more complex the task, the larger Fable 5’s lead over our other models.

    推荐理由:作者补充了基准之外的第一手使用感受,指出长难题求解能力跃升和安全护栏偏敏感,可作选型参考。

6月9日周二
6月6日周六
  1. Ahead of AI(RSS)49

    LLM 研究论文:2026 年清单(1 月至 5 月)

    Ahead of AI 发布 2026 年 1 月至 5 月的 LLM 研究论文精选清单,按架构与模型设计、高效训练与扩展、推理效率与 KV Cache、稀疏注意力与长上下文、推理与测试时计算、强化学习与 RLVR、智能体系统与工具使用、编码智能体与软件工程、扩散语言模型、模型评估与基准等类别整理。

6月4日周四
  1. vLLM 官方博客(RSS)67

    NVIDIA Nemotron 3 Ultra 在 vLLM 获 Day-0 支持

    NVIDIA 宣布 Nemotron 3 Ultra 在 vLLM 上获得 Day-0 支持。该开源模型采用混合 Transformer-Mamba 潜在 MoE 架构,总参数 550B、激活 55B,上下文最长 1M tokens,支持 NVFP4 与 BF16 精度,面向长时程自主智能体工作流。

    推荐理由:官方给出架构细节、显存配置和 vLLM 部署命令,读者可以直接照此在自有环境里跑通该模型。

6月2日周二
5月30日周六
5月27日周三
  1. Ethan Mollick:One Useful Thing(RSS)68

    Ethan Mollick 谈选择保持人类思考,AI 使用应有意而非默认依赖

    Ethan Mollick 撰文讨论 AI 无处不在的写作与认知投降风险,指出无脑依赖 AI 会削弱思考与学习。他引用土耳其约千名高中生的实验,用普通 ChatGPT 做作业的学生考试成绩反而低于无 AI 组;而台北十所高中近千名学生的 Python 课程中,AI 导师个性化出题让学生在无 AI 期末考试中高出 0.15 个标准差,相当于六到九个月额外学业。

    推荐理由:作者结合多项实验研究和自身写作经验,分析了默认依赖 AI 如何导致认知投降,并给出学习场景下避免捷径的具体做法。

5月26日周二
5月19日周二