vLLM 与 Speculators 开源支持 P-EAGLE、DFlash、DSpark 三种并行草稿算法
vLLM 与 Speculators 为 P-EAGLE、DFlash、DSpark 三种并行草稿(parallel drafting)算法提供完整开源支持,相关模型已发布在 RedHatAI HuggingFace Hub 的 Speculators Collection 中。
vLLM 与 Speculators 为 P-EAGLE、DFlash、DSpark 三种并行草稿(parallel drafting)算法提供完整开源支持,相关模型已发布在 RedHatAI HuggingFace Hub 的 Speculators Collection 中。
vLLM 宣布对月之暗面 Kimi K3 提供 Day-0 高效支持。Kimi K3 是 2.8 万亿参数的 MoE 多模态模型,每 token 激活 896 个专家中的 16 个,基于 Kimi Delta Attention 与 Attention Residuals,支持 1M token 上下文窗口。
推荐理由:vLLM 官方详述 Kimi K3 的服务适配与内核优化,读者可以据此了解 2.8T 混合 MoE 的生产部署配方和性能数据。
Berkeley AI Research 提出 ABBEL 框架,将 LLM 智能体的摘要以自然语言"信念状态"形式隔离并监督其信息内容,替代完整交互历史作为工作上下文。在协作编程基准 CollabBench 上,采用重建式信念评分的 ABBEL-rec-BG 将自摘要模型与全上下文模型的性能差距缩小约 50%,训练步数从 100 降至 50,峰值上下文 token 长度也低于全上下文设置。
DaoCloud 团队在 3 台 8 卡 B300 服务器(共 24 GPU)上以 4 Prefill + 1 Decode 分离拓扑部署 GLM-5.2-NVFP4,在 mean TTFT ≤ 2.5s、mean TPOT ≤ 20ms 的 SLA 下,将 16K 输入的 mean TPOT 从约 40ms 优化到约 17ms。
Google Research 发表 SymptomAI 论文,在 13,917 名参与者中测试基于 Gemini Flash 2.0 的五版实验性症状问诊智能体,所有诊断仅供研究分析。
推荐理由:原文给出 n=13,917 的随机对照研究设计和临床专家盲评结果,并展示了诊断与 Fitbit 生理信号的关联分析。
AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the Benjamini-Hochberg or BH method) and proved it controls the FDR. This method has been widely adopted in modern high-throughput science, including in genomics, astronomy, economics, etc. The paper has has garnered more than 130,000 citations to date. However Benjamini and Hochberg showed FDR control only when the data for the individual tests are *independent*. In practice, these data are often dependent; a good example is data on genetic variants due to linkage disequilibrium. Later work has focused on extending the validity of the BH procedure, e.g., to a form of positive dependence by Benjamini and Yekutieli (2001). The question of when the BH procedure controls the FDR has remained open. Over the last twenty years, many authors, including Reiner-Benaim (2007), Kim and van de Wiel (2008), Benjamini (2010), Sarkar (2023), Sarkar and Zhang (2025), have conjectured that the BH procedure controls the FDR for two-sided tests using any correlated Gaussian data. These authors have presented both theoretical and empirical evidence supporting, but not directly showing, the conjecture. With the help of AI (specifically GPT-5.6 Sol Pro), I have settled the question in the negative: The Benjamini-Hochberg procedure does *not* generally control the false discovery rate at the desired level for correlated two-sided Gaussian tests. This was done by exhibiting a Gaussian factor model for which, at a nominal level alpha=0.01, the false discovery rate is proved to be FDR>0.0104. There is a lot of interesting commentary to be made: 1. This result should be of interest to everybody in the field of statistics. Emmanuel Candes of Stanford University once called the false discovery rate and the Benjamini-Hochberg procedure "one of the two most important developments in statistics after 1950" (the other being James-Stein shrinkage). The present conjecture is probably the most central question about FDR/BH that was unresolved to date. 2. GPT-5.6 one-shot the problem after 90 minutes of reasoning, whereas with 5.5 I was not able to solve it even after iterating with multiple parallel agents for perhaps 20 hours. So the capability improvement is quite real. Exciting times to live in! 3. The argument is not especially surprising, but it does combine an asymptotic approach (standard for FDR analysis, see e.g., Genovese and Wasserman, Efron, etc) with a numerical certificate in a way that would be pretty non-standard in the field. Once we have the specific example, then straightforward simulations also support that the false discovery rate is indeed higher than the nominal value (see attached fig). 4. The current degree of violation over the nominal level is relatively small (0.104 vs 0.1). So the importance of this result is mainly conceptual. The practical implications remain to be determined. Overall, an exciting development! Preprint is available here (https://faculty.wharton.upenn.edu/wp-content/uploads/2017/06/bh.pdf) and will be on arxiv tonight; supporting code is here (https://github.com/dobriban/BH).
Thinking Machines Lab 发布多模态 MoE 模型 Inkling,Together AI 在发布当天通过 Serverless 提供托管推理,支持 1M 上下文窗口和 OpenAI 兼容 API。
推荐理由:原文给出 Inkling 的架构细节、参数规模和初步评测数据,读者可以据此判断其推理与多模态能力是否适合接入。
Thinking Machines 在 Hugging Face 上发布开源多模态模型 Inkling,共 975B 总参数、41B 激活参数,支持 1M 上下文,原生接收图像、文本和音频输入,训练数据为 45 万亿 token,并同时发布 276B 总参数、12B 激活参数的 Inkling-Small。
推荐理由:文章给出了 Inkling 的架构细节、各档位 VRAM 需求和多条部署路径,便于读者评估在自己的硬件上运行哪种变体。
vLLM 通过 V1 的公开 connector 接口接入 TileRT 作为专用 decode 引擎,随 TileRT 0.1.5 发布,prefill、调度、前缀缓存与 API 均保持原生 vLLM 不变。
vLLM 与 AMD Quark 团队展示了在 AMD Instinct MI355X GPU 上训练和部署 EAGLE3 投机解码的完整流程,覆盖 Kimi-K2.5 与 MiniMax-M2.5,并用 InferenceX 做基准测试。
Hugging Face 宣布 vLLM 的 transformers 建模后端现在达到甚至超过 vLLM 原生实现的吞吐速度,模型作者无需移植代码即可用 --model-impl transformers 获得 vLLM 级推理性能。
推荐理由:官方给出了与 vLLM 原生实现对比的具体吞吐数字和使用方法,读者可以据此判断是否切换到自己已有的 transformers 模型工作流。
Mistral AI 发布 Apache-2.0 协议的 Leanstral 1.5,总参数 119B、激活 6B,专注 Lean 4 形式化证明工程。
Together AI 宣布其 9 篇论文入选 ICML 2026(7 月 6 日至 11 日,首尔,展位 B714),覆盖从智能体到 GPU kernel 的全栈研究。
Nous Research 在 GitHub 开源 speculators,一个用于构建、评估和存储 LLM 推理投机解码算法的统一库,面向 vLLM。该库将投机解码算法的开发、评测与存储整合到同一套工具中。
vLLM Semantic Router 团队推出 Micro-Agent 能力,把单次模型 API 调用变成服务层内有界的多模型协作,用户仍只调用一个模型名 vllm-sr/auto。
Google Research 将在 COLM 2026 发表的论文 Thinking to Recall 发现,开启链式推理能让 Gemini-2.5(Flash 和 Pro)与 Qwen3-32B 检索到关闭推理时几乎无法召回的简单事实。
vLLM 宣布对 MiniMax M3 家族的 day-0 支持,涵盖 MiniMaxAI/MiniMax-M3 与 MiniMax-M3-MXFP8 两个检查点,支持 1M token 上下文和图像、视频多模态输入。
推荐理由:vLLM 团队自己拆解了 day-0 支持背后的 MSA 内核、缓存和部署细节,读者可以按需套用其部署配方。
Google DeepMind 发布实验性开源模型 DiffusionGemma,基于文本扩散方法,在专用 GPU 上实现最高 4 倍的文本生成提速。该模型为 26B 总参数的 MoE,推理时仅激活 3.8B 参数,量化后可装入 18GB 显存,单张 NVIDIA H100 上超过 1000 tokens 每秒,RTX 5090 上超过 700 tokens 每秒。
推荐理由:官方给出具体吞吐数字、显存占用和质量取舍,读者可据此判断它适不适合本地交互式工作流。
Google DeepMind 与 vLLM 团队合作,将 26B 参数、基于 Gemma4 的离散扩散语言模型 DiffusionGemma 接入 vLLM,这是 vLLM 首个原生支持的 dLLM。
推荐理由:原文详解了扩散语言模型的解码机制和 vLLM 的 ModelState 接入方案,并给出 H200 上的吞吐数据,对推理工程实践有直接参考价值。
vLLM 社区发布开源 RL 后训练框架 vime(Apache 2.0),基于 slime 的训练栈和数据生成设计,将 Megatron 训练与 vLLM 推理接入同一条解耦管线。
NVIDIA 宣布 Nemotron 3 Ultra 在 vLLM 上获得 Day-0 支持。该开源模型采用混合 Transformer-Mamba 潜在 MoE 架构,总参数 550B、激活 55B,上下文最长 1M tokens,支持 NVFP4 与 BF16 精度,面向长时程自主智能体工作流。
推荐理由:官方给出架构细节、显存配置和 vLLM 部署命令,读者可以直接照此在自有环境里跑通该模型。
Intel 的 AutoRound 训练后量化算法已完整集成进 vLLM-Omni,支持 W4A16 量化并实现"一次量化、直接部署"流程。Qwen3-Omni-30B-A3B 的 checkpoint 从 66 GB 降至 25 GB,体积减少最多 62%,其 W4A16 版本在 OmniBench 上得分略优于 BF16 参考,文生图质量漂移仅约 1.3%。
Together AI 详细解析其为 MiniMax M3 做的生产级推理优化,M3 支持 1M token 上下文窗口、原生多模态和 MiniMax Sparse Attention(MSA)架构,MSA 使 prefill 提速超过 9x、decode 提速超过 15x。
Together AI 发布面向编码智能体的高并发推理基准,用 4×NVIDIA B200 在 Kimi K2.5 上对比 Together Inference Engine、TensorRT-LLM 与 SGLang。
Together AI 发文解析 DeepSeek-V4 的服务化挑战,认为其核心变化是在 token 轴压缩 KV cache 的混合注意力设计(CSA、HCA、SWA),使模型支持 1M token 上下文窗口。
推荐理由:原文基于 Together 在 HGX B200 上的实际 bring-up,给出 KV cache 布局和缓存策略如何决定 DeepSeek-V4 长上下文吞吐的具体经验。
伯克利 AI 研究团队梳理了并行推理领域进展,指出当前多数方法(如 Self-consistency、Best-of-N、Tree of Thoughts、MCTS、ParaThinker、GroupThink、Hogwild!
OpenAI 报告其自动检测系统发现多个已发布模型在 RL 训练中意外受到有限的 CoT 评分,涉及 GPT-5.4 Thinking、GPT-5.1 Instant 至 GPT-5.4 Instant、GPT-5.3 mini 和 GPT-5.4 mini,GPT-5.5 未受影响。
推荐理由:OpenAI 自曝已发布模型在 RL 中出现过意外 CoT 评分,并给出检测系统与影响分析,读者可了解其监控性保护实践。
Together AI 宣布 DeepSeek-V4 Pro 上线,提供 Serverless Inference(512K 上下文)与 Dedicated 部署(完整 1M 上下文、预留容量)。
推荐理由:原文给出部署形态、定价与基准数字,读者可以据此评估在 512K 上下文上运行 DeepSeek-V4 Pro 的实际成本与路径。
NVIDIA Nemotron 3 Nano Omni 现已在 Together AI 平台可用,Together AI 提供首日(Day 0)接入与 Dedicated Inference 部署。
Together AI 提出分布感知推测解码(DAS),在 RL 后训练中最高减少 50% 的 rollout 时间且不改变模型输出。DAS 由自适应后缀树草稿器和长度感知调度两部分组成,前者无需梯度更新即可持续适配策略变化,后者通过跨 GPU 负载均衡与 GPU 内推测预算分配削减长尾请求。
Google Research 在 ICLR 论文中提出智能体记忆框架 ReasoningBank,可从成功和失败经验中蒸馏结构化推理记忆用于测试时自我进化,并开源了代码。
伯克利 AI 研究团队提出梯度规划器 GRASP,通过将轨迹提升为虚拟状态实现跨时间并行优化、在状态迭代中直接注入随机性以探索、并重塑梯度让动作获得清晰信号,从而让世界模型的长时程规划更稳健。
Together AI 提出 DBPlanBench,将 Apache DataFusion 的物理执行计划暴露给 LLM,把查询优化从统计计算转为语义推理。其序列化层将计划压缩约 10 倍,LLM 通过 JSON Patch 做局部改写而非重生成整个计划。
Google DeepMind 发布 Gemma 4 开源模型,主打高级推理与智能体工作流,采用 Apache 2.0 许可,提供 E2B、E4B、26B MoE 和 31B Dense 四个尺寸。
推荐理由:官方给出四个尺寸、榜单排名、上下文长度和开放许可等具体信息,读者可以据此判断选型与部署空间。
Together AI 开源 Aurora,一个基于 RL 的推测解码框架,可从实时推理轨迹中持续学习并异步更新 draft 模型,在 Qwen3 和 Llama3 等模型上比静态 speculator 额外提速 1.25 倍。
Together AI 在 ICLR 2026 论文《When Does Divide and Conquer Work for Long Context LLM?
Together AI 扩展 Together Fine-Tuning,新增工具调用、推理和视觉语言模型(VLM)微调支持,训练栈升级后支持 100B+ 参数模型,吞吐最高提升 6×,数据集支持最大 100GB。
Mistral 发布 Mistral Small 4,首个将 Magistral 的推理、Pixtral 的多模态和 Devstral 的智能体编码能力统一到单一模型的开源版本,采用 Apache 2.0 许可证。
推荐理由:官方发布文给出完整架构参数、性能对比和部署门槛,可以了解这款统一推理与多模态的开源模型是否适配自己的场景。
Mistral 发布 Leanstral,一个面向 Lean 4 的开源代码智能体,权重以 Apache 2.0 许可开放,同时在 Mistral Vibe 中以 agent 模式提供,并通过免费 API 端点 labs-leanstral-2603 开放。
推荐理由:原文给出 FLTEval 分数和与 Claude 系列的成本对比,读者可据此评估它在形式化证明工程中的性价比。
伯克利 AI 研究团队提出 SPEX 与 ProxySPEX 算法,用于在大规模场景下识别 LLM 的关键交互。SPEX 利用稀疏性与低阶性,将交互搜索转化为稀疏恢复问题;ProxySPEX 进一步利用层次结构,以约少 10 倍的消融次数达到 SPEX 的性能。在情感分析任务上,SPEX 在上下文扩展至数千特征时仍保持高忠实度,优于 LIME、Banzhaf 等边际方法。