跳到正文

#数据/训练

今日 93 条
9月24日周四
  1. NVIDIA Technical Blog(开发者技术博客 · RSS)21

    NVIDIA 如何为生物基础模型实现高效 MoE 训练

    NVIDIA 技术博客解析了生物基础模型的高效 MoE 训练方案:相比每个 token 都要过全部层的稠密 Transformer,MoE 用多个专家子网络、每个 token 只激活其中一小部分,从而降低训练与推理算力开销。文章指出,随着语言模型规模增长,稠密架构的扩展成本越来越高。

  2. Anthropic48

    Claude 在噬菌体 DNA 中发现了一个此前未知的酶系统,其基因旁有一段类似 CRISPR 的重复 DNA 序列。目前尚不清楚该系统功能,但仅有少数已知系统具备类似特征,且它们都能剪切、复制和粘贴 DNA。历史上此类可编程系统的发现曾推动医学变革,CRISPR 就是如今基因药物的基础,但还需更多研究才能确定该系统能否有类似用途。

9月23日周三
  1. NVIDIA Blog(RSS)12

    NVIDIA 验证工程师 Sakeena Fiza 如何让硬件大规模成功

    NVIDIA 验证工程师 Sakeena Fiza 在数据中心系统工程实验室负责新产品量产前的硬件验证,她回忆 NVIDIA Rubin GPU 首次在系统级被识别时全场欢呼,那是全球首次 Rubin GPU 在系统层面完成枚举。她将验证工作比作破案,目标是在客户之前发现问题,单块板卡可能含数万个组件,一个机架接近 50 万个。

  2. Ant Ling46

    感谢 @ValsAI 的高水准评测!“flash”这个词现在有点“误导”了。凭借 124B 总参数和 5.1B 激活,Ling-3.0-flash-fin 是一款高智能密度的“flash lite”。趁免费 API 还在,尽情享用吧。我们还有可用于本地 AI 的 fp4 量化 😛

    引用Vals AI@ValsAI

    Ant Group’s Ling 3.0 Flash Fin is a finance-specialized open-weight model that delivers strong financial analysis at budget-model pricing. On Finance Agent v2, it scores 54.9% at just $0.045 per task.

  3. StepFun55

    阶跃星辰(StepFun)宣布开源内部使用的 LLM 数据标注与模型检查工具 onPanda,工作流为找到错误、修正 token、让模型继续生成。数据标注方面,标注时间中位数比人工后编辑降低 52%,SFT 与偏好数据可在同一流程完成(ΔPPL <1%),支持 token 级正负样本监督及图像、音频、视频上的 agent 轨迹标注。

    引用Lei Yang@diyerxx

    I spent two years building this interactive tool to let you steer LLMs and agents at the token level. Introducing onPanda — a web app for token visualization & control, model inspection, data annotation, and more. Try it online (works on mobile): https://onpanda.diyer22.com/

9月22日周二
  1. karminski-牙医40

    卧槽5-10T???

    引用Max For AI@MaxForAI

    🚨Qwen4家族首次曝光!! 刚刚,在2026年云栖大会的开幕式上,新任@Alibaba_Qwen LLM负责人刘大一恒官宣了即将到来的Qwen4家族! 包含Qwen4-Max Qwen4-Flash&amp;Qwen4-Plus 还有Qwen4-27B!!! 未来Qwen会训5-10T的模型

  2. Latent Space(RSS)79

    Xiaomi MiMo-V2.6-Pro 1T-A42B 登顶开源权重模型,训练仅花费约 $3M

    Latent Space AINews 汇总 2026/9/19-9/21 AI 动态,核心是 Xiaomi 发布 MiMo-V2.6-Pro(1.02T 总参数/42B 激活,MIT 许可),以 Artificial Analysis Intelligence Index 46 分成为新的开源权重榜首,成本为 $0.435/M 输入、$0.87/M 输出 token。

    推荐理由:除发布信息外还汇总了 RL 成本与训练细节,读者可以看到开源权重模型追赶闭源的具体路径。

9月21日周一
  1. MIT Technology Review · AI36

    MIT Technology Review 如何绘制首张美国边境“虚拟墙”沿线死亡地图

    MIT Technology Review 与 Times of San Diego 合作,用 15 个月完成首张美国边境监控塔附近移民死亡的综合地图与分析,数据回溯至 2015 年。团队向得州 17 个县警长办公室申请记录,收到超 4000 页文件,并对 Kenedy、Webb、Hidalgo 三县的记录调用 Anthropic Claude API 提取遗骸发现坐标后人工核验。

9月20日周日
  1. elsewhere:文章(RSS)60

    峰瑞李丰:全球流动性见顶后,AI周期进入后半段的投资逻辑

    峰瑞资本李丰撰文分析,认为2026年三季度全球流动性接近见顶,美元主导的资本市场进入存量博弈尾部,AI产业周期进入后半段。文章回顾2020年天量流动性如何催生本轮AI热潮,列举科技巨头资本开支转折的五个信号(如Alphabet二季度自由现金流转负59亿美元),提出投资重心应从讲大故事转向能靠AI赚钱的方向,如AI+应用、生物医疗与AI交叉及SaaS的AI化。

9月19日周六
9月18日周五
9月16日周三
  1. MIT News(RSS)23

    MIT 政治学者 Naoki Egami 如何用统计方法研究社会测量与 AI 工具误差

    MIT 政治学副教授 Naoki Egami 专注研究方法论,尤其研究社会科学的“外部有效性”,即特定研究结论能否推广到其他情境。他早在 ChatGPT 引发 AI 热潮之前就开始研究 AI 工具引入研究后产生的误差,以及如何系统识别并校正这些误差。Egami 2020 年获普林斯顿大学博士学位,2025 年加入 MIT 政治学系。

  2. Apple Machine Learning Research(RSS)38

    Glyph:面向企业数据目录列描述与敏感本体标注的多策略智能体系统

    Apple 研究团队提出 Glyph,一个将列描述生成与列类型标注建模为有状态图编排的多智能体 LLM 生产系统。其 Descriptor 通过推理-行动工具循环从企业 GitHub 按需检索管道源码来支撑生成,Tagger 并行运行描述、业务线正则与元数据三种策略,并用 RRF 融合排序结果,从 275 叶节点的数据分类本体中打标。

  3. Jeff Dean42

    令人振奋的成果,@LiamFedus!祝贺 Periodic Labs 的整个团队!

    引用Liam Fedus@LiamFedus

    We built high-throughput materials labs in Menlo Park to create a loop between experiments and models. The labs generate fresh data, the models learn from it, and then help us decide what to try next. Using only 1,300 H200s, plus months of our experimental data, we mid-trained and RL’d an open-source model to surpass GPT-6 Astra on our analysis benchmark. We call it Neon. This is real footage from our lab. We’re focusing first on hard problems in materials science, including superconductors, magnets, and semiconductor materials. Read our blog posts below.

  4. Google Research29

    Google Research 提出 Retrieve-for-Train:用 RL 编译扩散模型绕过推理瓶颈,加速复杂 AI 搜索

    Google Research 提出 Retrieve-for-Train 框架,通过离线强化学习发现奖励对齐的查询扇出并编译为监督信号,再蒸馏进一个 53.9M 参数的扩散检索器,实现推理时单次非自回归的查询扇出。该方法在 Gemma3-4B 和 Qwen3-4B 上微调扇出语言模型,用集合级属性奖励评估整组结果,无需人工标注,也无需推理时的 CoT 思考 token。

9月15日周二
  1. NVIDIA Technical Blog(开发者技术博客 · RSS)33

    NVIDIA Transformer Engine 如何加速 JAX 中的 Dropless MoE 训练

    NVIDIA 技术博客介绍如何用 NVIDIA Transformer Engine 在 JAX 中加速 Dropless MoE 训练。MoE 通过条件计算实现高效训练,DeepSeek、Qwen、Mixtral 等模型以远低于稠密模型的训练算力达到或超越其性能。文章针对传统 MoE 依赖共享稠密 FFN 的做法,给出 Dropless 训练路径。

9月14日周一
9月12日周六
  1. Thinking Machines42

    我们自己的 @johnschulman2 与 Dwarkesh 对话,讨论随着模型不断进步和自我改进,人类判断力在哪些方面仍然重要:教它们处理混乱的现实世界任务,用品味判断什么在长期内有效,以及最重要的——明确我们真正想要什么。

    引用Dwarkesh Patel@dwarkesh_sp

    New episode with @johnschulman2, @oneill_c and @BerenMillidge. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. 0:00:00 – Steelmanning the case against RSI 0:18:39 – What’s driving the Chinese labs’ progress 0:28:06 – How will automated AI researchers be trained 0:33:51 – Will long-horizon RL elicit AGI? 0:45:24 – The sim-to-real gap 1:00:33 – How much progress is explained by data? 1:18:03 – Why is RL working so well? 1:24:54 – Move 37 and entropy collapse 1:28:31 – Rapid-fire timelines

  2. Dwarkesh Patel:Podcast & Blog(RSS)61

    Dwarkesh 对谈 John Schulman、Beren Millidge 与 Charlie O'Neill:AI 研究者激辩递归自我改进还有多远

    Dwarkesh Patel 邀请 Zyphra CTO Beren Millidge、Thinking Machines 首席科学家 John Schulman 和 Baseten 模型训练负责人 Charlie O'Neill 对谈递归自我改进(RSI)何时到来。

    推荐理由:三位一线研究者围绕递归自我改进给出了各自不同的技术瓶颈判断,涵盖蒸馏、sim-to-real 与持续学习等具体分歧。

9月11日周五
9月10日周四
9月9日周三
  1. Mark Chen38

    两件事要区分: 在 Navier Stokes 工作中,有任何人类或智能体查看过用户数据吗?没有。 我们是否以整体方式使用用户反馈和去标识化数据来改进 ChatGPT 和 Codex?是的。每一家 LLM 公司都是如此。

    引用levent@__alpoge__

    “we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” i mean props to them for straight coming clean. (so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan) so i’ll now give a bit on my thinking here. i actually woulda been pumped to collaborate on this, there are a lot of people at oai i like (ok, clearly some were indirectly dicks to me because of being part of the whole situation, but im a big boy, i still like them), idgaf about authorship on that step anyway, coulda been me Tristan and every fte at oai for all i care (on that Tristan would disagree:p). but on hearing the loud convo in the hallway, especially the part where a millennium prize was offered if i’d just be removed from the paper, it was kinda clear the die had been cast and things were locked. pretty wacky, unstrategic, and unnecessary, since on my side things were mostly me and claude having a good time yoloing random stuff in the corner rather than anything institutional. i also like the idea of the labs cooperating, and even better on scientific progress. it’s a shame!

  2. Dwarkesh Patel:Podcast & Blog(RSS)56

    Dwarkesh Patel 分析:2019-2025 年预训练进展主要来自数据改进

    Dwarkesh Patel 通过训练 2019-2025 年各年度代表性模型配方与数据语料的组合(最高 1e19 FLOPs,用 OLMES 评估)发现,数据改进带来 12.0x 计算效率提升,模型改进为 3.7x,数据贡献约为模型的 3.24 倍;模型与数据收益基本相互独立,88% 的 OLMES 分数方差可由二者的加性效应解释。

9月8日周二
  1. Google DeepMind:Blog(RSS)80

    Google DeepMind 发布 AlphaGenome Atlas,覆盖人类基因组 90 亿个单核苷酸变异预测

    Google DeepMind 发布 AlphaGenome Atlas,一个包含人类基因组全部约 90 亿个单核苷酸变异效应预测的平台,规模达 1PB,是 AlphaFold Database 的 30 倍以上。

    推荐理由:AlphaGenome Atlas 把 90 亿个单核苷酸变异的预测结果做成可检索资源,读者可了解其数据规模与在罕见病研究中的验证案例。

9月5日周六
  1. a16z:News(RSS)43

    a16z 投资 Gimlet Labs:打造首个多芯片推理云

    a16z 宣布投资 Gimlet Labs,后者正在构建首个多芯片推理云,可在同一功耗范围内为前沿模型带来最高 10 倍吞吐与交互性提升。Gimlet 通过编译器与运行时把不同模型和工具调度到 GPU、CPU 及专用加速器上,并将编排延伸至数据中心层面,对开发者只暴露单一推理 API。其客户已包括一家前沿实验室和一家超大规模云厂商。