NVIDIA 如何为生物基础模型实现高效 MoE 训练
NVIDIA 技术博客解析了生物基础模型的高效 MoE 训练方案:相比每个 token 都要过全部层的稠密 Transformer,MoE 用多个专家子网络、每个 token 只激活其中一小部分,从而降低训练与推理算力开销。文章指出,随着语言模型规模增长,稠密架构的扩展成本越来越高。
NVIDIA 技术博客解析了生物基础模型的高效 MoE 训练方案:相比每个 token 都要过全部层的稠密 Transformer,MoE 用多个专家子网络、每个 token 只激活其中一小部分,从而降低训练与推理算力开销。文章指出,随着语言模型规模增长,稠密架构的扩展成本越来越高。
我们与 @GoogleDeepMind、@emblebi 及研究伙伴合作,将 2,800 多种病毒的 AI 预测蛋白质复合物结构公开开放。 这为科学家提前应对潜在疫情提供了先机。
NVIDIA 联合 Google DeepMind、EMBL-EBI 等机构,通过 AlphaFold Database 开放了 2800 多种病毒的蛋白复合物预测 3D 结构,供全球科学家免费使用。
推荐理由:读者可了解开放病毒蛋白结构数据集如何降低结构预测门槛,以及配套开源流程的复用方式。
NVIDIA 验证工程师 Sakeena Fiza 在数据中心系统工程实验室负责新产品量产前的硬件验证,她回忆 NVIDIA Rubin GPU 首次在系统级被识别时全场欢呼,那是全球首次 Rubin GPU 在系统层面完成枚举。她将验证工作比作破案,目标是在客户之前发现问题,单块板卡可能含数万个组件,一个机架接近 50 万个。
Ant Group’s Ling 3.0 Flash Fin is a finance-specialized open-weight model that delivers strong financial analysis at budget-model pricing. On Finance Agent v2, it scores 54.9% at just $0.045 per task.
I spent two years building this interactive tool to let you steer LLMs and agents at the token level. Introducing onPanda — a web app for token visualization & control, model inspection, data annotation, and more. Try it online (works on mobile): https://onpanda.diyer22.com/
🚨Qwen4家族首次曝光!! 刚刚,在2026年云栖大会的开幕式上,新任@Alibaba_Qwen LLM负责人刘大一恒官宣了即将到来的Qwen4家族! 包含Qwen4-Max Qwen4-Flash&Qwen4-Plus 还有Qwen4-27B!!! 未来Qwen会训5-10T的模型
Latent Space AINews 汇总 2026/9/19-9/21 AI 动态,核心是 Xiaomi 发布 MiMo-V2.6-Pro(1.02T 总参数/42B 激活,MIT 许可),以 Artificial Analysis Intelligence Index 46 分成为新的开源权重榜首,成本为 $0.435/M 输入、$0.87/M 输出 token。
推荐理由:除发布信息外还汇总了 RL 成本与训练细节,读者可以看到开源权重模型追赶闭源的具体路径。
Latent Space 播客访谈 TypeSafe AI CEO Diogo Almeida,介绍其新发布的 Jev 模型,定位为面向软件而非聊天的 System 1 可编程模型,优化智能与成本之比。
微软研究院在 Nature 发表逆合成模型 RetroChimera,并开源其实现与权重。该模型用学习式集成策略融合 Transformer 模型 R-SMILES 2 与 GNN 模型 NeuralLoc 的排序预测。
MIT Technology Review 与 Times of San Diego 合作,用 15 个月完成首张美国边境监控塔附近移民死亡的综合地图与分析,数据回溯至 2015 年。团队向得州 17 个县警长办公室申请记录,收到超 4000 页文件,并对 Kenedy、Webb、Hidalgo 三县的记录调用 Anthropic Claude API 提取遗骸发现坐标后人工核验。
小米 MiMo 在 GitHub 新建仓库 verl,其 HybridFlow 是一个灵活高效的强化学习后训练框架。该仓库定位为 RL post-training 框架,强调灵活性与效率,目前原文未披露参数规模、benchmark 分数或开源许可等细节。
百曜科技联合《麻省理工科技评论》中国团队发布《AI 虚拟细胞(AIVC)技术趋势、产业生态与应用前景研究报告》,将 AIVC 定义为 AI 时代生命科学的新型基础设施。
峰瑞资本李丰撰文分析,认为2026年三季度全球流动性接近见顶,美元主导的资本市场进入存量博弈尾部,AI产业周期进入后半段。文章回顾2020年天量流动性如何催生本轮AI热潮,列举科技巨头资本开支转折的五个信号(如Alphabet二季度自由现金流转负59亿美元),提出投资重心应从讲大故事转向能靠AI赚钱的方向,如AI+应用、生物医疗与AI交叉及SaaS的AI化。
Google Research 发布 MilleMiglia,一个用 C++ 编写的实例生成器,用于为中段物流(middle-mile)配送问题生成真实且保护隐私的基准数据,源码与文档已在 GitHub 开放。
SemiAnalysis 研究了 DeepSeek-V4.1-Flash 的 Engram 多 token 查找机制,其表可卸载到主机 DRAM,同质量下降低 HBM 需求。
MIT 政治学副教授 Naoki Egami 专注研究方法论,尤其研究社会科学的“外部有效性”,即特定研究结论能否推广到其他情境。他早在 ChatGPT 引发 AI 热潮之前就开始研究 AI 工具引入研究后产生的误差,以及如何系统识别并校正这些误差。Egami 2020 年获普林斯顿大学博士学位,2025 年加入 MIT 政治学系。
Apple 研究者提出 DACA-GRPO,一种可插拔的 GRPO 训练增强方法,用于扩散语言模型强化学习。它通过 Denoising Progress Scores 提取逐 token 重要性权重(无额外前向开销),并用 Stratified Masking Likelihood 降低 mean-field 似然偏差。
Apple 研究团队提出 Glyph,一个将列描述生成与列类型标注建模为有状态图编排的多智能体 LLM 生产系统。其 Descriptor 通过推理-行动工具循环从企业 GitHub 按需检索管道源码来支撑生成,Tagger 并行运行描述、业务线正则与元数据三种策略,并用 RRF 融合排序结果,从 275 叶节点的数据分类本体中打标。
令人振奋的成果,@LiamFedus!祝贺 Periodic Labs 的整个团队!
We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.
Latent Space 访谈 Good Start Labs CEO Alex Duffy,探讨用 Diplomacy、1830 等游戏训练 AI 模型能否让技能迁移到真实工作。
Google Research 提出 Retrieve-for-Train 框架,通过离线强化学习发现奖励对齐的查询扇出并编译为监督信号,再蒸馏进一个 53.9M 参数的扩散检索器,实现推理时单次非自回归的查询扇出。该方法在 Gemma3-4B 和 Qwen3-4B 上微调扇出语言模型,用集合级属性奖励评估整组结果,无需人工标注,也无需推理时的 CoT 思考 token。
NVIDIA 详解 NVLink 6 如何为大规模 AI 工厂提供多层弹性。在超大规模 AI 训练中,集群内每块 GPU 每秒需在数千次集合通信操作中同步梯度;推理阶段非计划停机则直接减少请求处理总量,限制营收。
NVIDIA 技术博客介绍如何用 NVIDIA Transformer Engine 在 JAX 中加速 Dropless MoE 训练。MoE 通过条件计算实现高效训练,DeepSeek、Qwen、Mixtral 等模型以远低于稠密模型的训练算力达到或超越其性能。文章针对传统 MoE 依赖共享稠密 FFN 的做法,给出 Dropless 训练路径。
vime 联合 RL-Kernel 在 AMD Instinct MI300X 上实现训练与 rollout 的逐位数值一致性,8× MI300X 跑 Qwen3-8B GRPO 实验连续 200 步 mismatch_count = 0、max_abs_diff = 0。
New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines
Dwarkesh Patel 邀请 Zyphra CTO Beren Millidge、Thinking Machines 首席科学家 John Schulman 和 Baseten 模型训练负责人 Charlie O'Neill 对谈递归自我改进(RSI)何时到来。
推荐理由:三位一线研究者围绕递归自我改进给出了各自不同的技术瓶颈判断,涵盖蒸馏、sim-to-real 与持续学习等具体分歧。
Together AI 扩展 Together Fine-Tuning 服务,新增 GLM 5.3、Kimi K2.7、DeepSeek-V4-Flash、Qwen 3.8-27B、Gemma 4 等开源权重模型支持,并通过 API、CLI 和 UI 提供逐步训练指标实时追踪。
Google Research 提出 ToolGrad,一种“先答案后问题”的工具调用数据生成范式:先迭代构建可验证的 API 调用链,再反推用户提示词,替代 ToolBench、ToolACE 等基于 DFS 试错的低效方案。
哦,原来这就是我们绘制出雄性果蝇全部 166,000 个神经元的原因。来看看这项社区大工程,展示这些微小果蝇大脑究竟有多大能耐 🪰🧵
编码器回归了,宝贝 (引用推文 @scaling01:这到底是什么外星架构)
what in the alien architecture is this
SemiAnalysis 报告称其能源模型已追踪到 75GW 表后 AI 算力的确定性订单,仅 2026 年 Q2 就新增约 20GW,微软年内签署超 5GW,OpenAI 将在得州 Shackelford County 启用 1.4GW 离网园区。
“we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” i mean props to them for straight coming clean. (so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan) so i’ll now give a bit on my thinking here. i actually woulda been pumped to collaborate on this, there are a lot of people at oai i like (ok, clearly some were indirectly dicks to me because of being part of the whole situation, but im a big boy, i still like them), idgaf about authorship on that step anyway, coulda been me Tristan and every fte at oai for all i care (on that Tristan would disagree:p). but on hearing the loud convo in the hallway, especially the part where a millennium prize was offered if i’d just be removed from the paper, it was kinda clear the die had been cast and things were locked. pretty wacky, unstrategic, and unnecessary, since on my side things were mostly me and claude having a good time yoloing random stuff in the corner rather than anything institutional. i also like the idea of the labs cooperating, and even better on scientific progress. it’s a shame!
Dwarkesh Patel 通过训练 2019-2025 年各年度代表性模型配方与数据语料的组合(最高 1e19 FLOPs,用 OLMES 评估)发现,数据改进带来 12.0x 计算效率提升,模型改进为 3.7x,数据贡献约为模型的 3.24 倍;模型与数据收益基本相互独立,88% 的 OLMES 分数方差可由二者的加性效应解释。
Google DeepMind 发布 AlphaGenome Atlas,一个包含人类基因组全部约 90 亿个单核苷酸变异效应预测的平台,规模达 1PB,是 AlphaFold Database 的 30 倍以上。
推荐理由:AlphaGenome Atlas 把 90 亿个单核苷酸变异的预测结果做成可检索资源,读者可了解其数据规模与在罕见病研究中的验证案例。
NVIDIA 团队用 NemoClaw 构建了一个记忆驱动的"幕僚长"智能体,通过名为 self model 的人类可读知识层保存智能体记忆。该智能体面向企业工作中随时间变化的邮件、决策、项目和待办事项,避免每次启动都需重建上下文。
a16z 宣布投资 Gimlet Labs,后者正在构建首个多芯片推理云,可在同一功耗范围内为前沿模型带来最高 10 倍吞吐与交互性提升。Gimlet 通过编译器与运行时把不同模型和工具调度到 GPU、CPU 及专用加速器上,并将编排延伸至数据中心层面,对开发者只暴露单一推理 API。其客户已包括一家前沿实验室和一家超大规模云厂商。