We spent >20B tokens throwing @openai's Astra at every AI Engineering task we could think of, beyond cute Blender demos and fun games. https://latent.space/p/astra Here's everything Astra can do, and do so at <$6 an hour (serious): - choose and train models - label data (both helping you label and then using your labels for active learning) - keep pipelines saturated - instrument and read logs - deploy and debug entire systems in one shot - fan out and command and eval subagents (including agents running other models) - keep coherence over billions of tokens of a single agent thread. more to come on @swyx's coverage of the Fable- and Astra-class of 2026!
#数据/训练
#数据/训练
今日 93 条
swyx@swyxAI 评分6565引用Latent.Space@latentspacepodGoogle ResearchAI 评分2727 Google Research 研究跨人群多基因风险评分的迁移学习
Google Research 用 UK Biobank 的欧洲人群数据向 Biobank Japan 近 20 万日本人样本做迁移学习,评估 8 种临床性状的多基因风险评分(PRS)跨人群表现。
Google Research精选AI 评分6262 Google 与 HHMI Janelia 发布完整雄性果蝇脑连接组图谱
Google Research 与 HHMI Janelia 及剑桥等机构合作,在 Cell 发表论文,发布完整雄性果蝇脑与中枢神经系统连接组图谱,包含超过 166,000 个神经元和 1.25 亿个突触连接,是迄今按神经元数量计最大的脑图谱。
推荐理由:读者可了解 AI 重建如何把电子显微镜切片拼成完整脑图谱,以及这一资源对神经科学研究的用途。
Google DeepMind:Blog(RSS)精选AI 评分7474 Google DeepMind 发布 WeatherNext 3 全球天气 AI 模型
Google DeepMind 与 Google Research 发布 WeatherNext 3 全球天气 AI 模型,直接学习实时地球静止卫星数据,每小时生成一次预报,地表变量分辨率达 5 公里,整体比 WeatherNext 2 的 25 公里网格清晰约 5 倍。
推荐理由:对比前代的分辨率、更新频率与降水评分提升,可了解实时卫星数据如何改变全球天气预报的精度边界。
Hugging Face:Blog(RSS)AI 评分5959 用 TRL 和 OpenEnv 复现训练编码模型画水彩画
Hugging Face 博客发布复现实践:作者用 TRL 和 OpenEnv 开源复现 Surya Narreddi 用 RL 让 Qwen/Qwen3.5-35B-A3B 写 p5.brush JavaScript 画水彩的方法,参考池、RL 环境、训练脚本和模型全部公开。
Meta Engineering Blog(RSS)AI 评分6363 Meta 工程博客详解组织级第二大脑:让 AI 智能体从领域专家身上学习
Meta 工程团队构建了一个面向合规领域的 AI 智能体,作为组织级“第二大脑”,通过结构化知识文件与可组合 recipes 分离知识与推理,并以自动化自我改进闭环将专家反馈编译为经回归测试的更新,无需模型重训。
SemiAnalysis 长文 RSS(RSS)AI 评分6464 SemiAnalysis 深度分析韩国万亿美元主权 AI 投资:Nvidia 受益、Hynix 承压
SemiAnalysis 深度分析韩国主权 AI 战略:政府以锦标赛制推进“独立 AI 基础模型”项目,从 15 个联合体中选出 Naver Cloud、LG AI Research、SK Telecom、NC AI、Upstage 五队,首轮后淘汰 NC AI 并因使用阿里 Qwen 视觉与音频编码器取消 Naver 资格,递补 Motif Technologies。
Google ResearchAI 评分4242 Google 用深度学习从太空绘制全球甲烷排放地图
Google 提出 MAPL-EMIT 深度学习框架,基于 EMIT 高光谱辐射数据自动检测、预测增强并定位全球甲烷羽流,在专家标注羽流上召回率达 84%。该模型采用 Swin-S 视觉 Transformer,同时完成增强量化、羽流分割与源定位三项任务,训练数据为注入真实 EMIT 场景的 360 万个合成甲烷羽流。
Microsoft Research 博客(RSS)AI 评分5252 微软发布 GigaPath-Flash 与 GigaTIME-Flash 高效病理基础模型
微软研究院联合华盛顿大学和 Providence 发布 GigaPath-Flash 与 GigaTIME-Flash,两者均以 Apache 2.0 许可在 Hugging Face 开放权重。
Jensen Huang@JensenHuangAI 评分4040引用Gavin Baker@GavinSBakerRegret the tone of my post on data centers yesterday. What I should have said: There were reasonable concerns about data centers 18ish months ago: water, taxes, jobs, electricity prices, the environment and what they would do to small towns. Well-structured data center projects have largely addressed these concerns today and we should be celebrating this. On balance, data centers are awesome for America in every way. On water: U.S. data centers use a fraction of what golf courses use. A lot of the numbers from 18 months ago were off by over 1000x. Newer data centers use closed-loop systems or recycled water. Should be required by every town approving a data center project. On taxes: looking only at sales-tax exemptions, as Ronan Farrow did, is the wrong way to evaluate this. Data centers pay significant property taxes. Loudoun County, which is the wealthiest county in America, now collects on the order of $1 billion a year from data centers. In Quincy, WA, data centers are more than half the property-tax roll. Over time, property taxes can go to zero while government spending increases in these towns. On jobs: this has been unambiguously awesome for blue collar Americans. Demand for electricians, plumbers, welders, HVAC techs, and contractors has gone vertical, and it is not a one-time construction job. These buildings get upgraded and expanded over time. That is why the building trades are fighting for them, and why some unions are now treating opposition to data centers as a reason not to endorse politicians. On power: the original fear was that households would pay for the incremental electricity demand in the form of higher prices. That is why the ratepayer-protection deals and the new large-load tariffs exist. The right structure is: the data center brings or pays for new generation and signs a contract long enough that existing customers are protected. Where that is happening, utilities are cutting or freezing residential rates and saying so on the record. Where it is not, people are right to object. Electricity prices are going down *today* in a number of large states because of data centers. On the environment: data centers overwhelming use natural gas today, which is the cleanest power source outside of nuclear, solar and wind. And the companies that are building the data centers are committed to carbon neutrality such that an equivalent amount of solar will likely be built. Maybe more importantly, the data centers need batteries to function effectively and these batteries can also sell energy back into the grid (which recently prevented blackouts in Texas). Over time, data centers will run on solar plus batteries. On the towns: Poverty in Quincy, WA fell from 29% to 6%. Data center taxes paid for a new high school, a hospital, a library, police and fire stations. This is happening in many left for dead former mill and farm towns that had no other bidder for the land. Data centers are actually reindustrializing parts of America and creating the kind of working-class jobs both parties have spent decades claiming to support. That should not be a partisan issue. Data centers can and should be awesome for America and they increasingly, overwhelmingly are. Supporting the outsourcing of data centers to China will likely age just as well as support for the outsourcing of high quality, blue collar manufacturing jobs to China has aged. When the facts change, I change my mind. I hope that reasonable people who had good faith reasons to oppose data centers at least consider updating their beliefs given the change in the facts over the last 18 months. This really matters for America. I will say I also think the idea of making data centers beautiful is a good one that has yet to be implemented. Data centers should be just as beautiful as Grand Central Station. We can learn a lot from the railroad buildout. Neoclassical revival ftw. Might write up open-weight AI tomorrow as this is equally essential to America.
elsewhere:文章(RSS)AI 评分3636 Pyromind 创始人 Kevin Ding 谈 AI 下半场:Agent 蜂群与 AutoRL 而非单一超级模型
Pyromind 创始人兼 CEO Kevin Ding 在播客中提出,AI 终局更像 Agent 蜂群而非超级基础模型一统天下,公司押注 AutoRL 而非仅做 RL as a Service。
Meituan LongCat@Meituan_LongCatAI 评分5959美团 LongCat 发布 AutoResearchEval 评测,评估 7 个前沿模型在 36 项需要持续实验的 AI 研发任务上的表现,共覆盖 756 条轨迹。
Thinking Machines@thinkymachinesAI 评分4545引用Tinker@tinkerapiLLMs with scaffolds have lagged on text-to-SQL, a task that relies on human judgment. By folding expert judgment into every part of RLVR on Tinker, @maxYuxuanZhu and @ddkang (UIUC and Bridgwater) trained the first text-to-SQL model to beat the human mark. https://thinkingmachines.ai/news/putting-task-expertise-into-rl
MIT News(RSS)AI 评分3535 MIT 团队提出 PottsMPNN:不再以还原天然序列衡量蛋白质设计
MIT 生物学系团队开发出机器学习框架 PottsMPNN,通过引入支配蛋白质结构与稳定性的物理原理并建模氨基酸两两相互作用,提升序列生成与突变稳定性预测能力,成果发表于 PNAS。研究者指出,长期以来以能否复现进化选出的天然序列作为成功标准并非蛋白质设计的最佳指标,PottsMPNN 在减少对天然序列依赖的同时,结构兼容性与能量预测反而改善,可设计出序列不类似任何天然蛋白的结构可行蛋白。
MIT News(RSS)AI 评分4343 MIT 提出 CrysVCD 框架:让 AI 生成的材料更稳定、更贴近真实应用
MIT 研究人员提出 CrysVCD 框架,在材料生成前用语言模型约束价电子规则,使常用材料模型在近 70% 的生成结果中达到高晶格动力学稳定性。该方法比生成后再筛选的方案效率高一个数量级,微调后生成的晶体材料机械稳定性达 68%、亚稳性达 85%,并可定向生成高热导率、高介电常数等材料。
Hugging Face:Blog(RSS)精选AI 评分6666 IBM 发布 Granite 4.2 推理模型系列并详解构建过程
IBM Granite 团队发布 Granite 4.2 推理模型系列,包含 3B、8B、30B 三个 dense 版本,基于 Granite-4.1 基座(约 15T tokens 预训练,上下文扩展到 512K),经 SFT 和多阶段 GRPO 强化学习训练,8B 和 30B 额外经历 SWE、终端、搜索三类真实环境 agentic RL。
推荐理由:官方完整披露了从预训练、SFT 到多阶段 GRPO 强化学习的训练细节,读者可以据此了解推理模型的完整构建流程。
Hugging Face:Blog(RSS)AI 评分5555 Multiverse Computing 提出 Quantization-Aware Healing:60B 4-bit 模型在 9 项基准中 7 项反超全精度原版
Multiverse Computing 发布论文 Quantization-Aware Healing(QAH),将 GPT-OSS 120B 压缩到 60B 参数并量化为 MXFP4 后,直接从压缩前的原始模型做 KL 蒸馏。
MIT News(RSS)AI 评分4444 MIT 工程师开发 η-learning:无需极端事件数据即可生成极端天气场景
MIT 工程师提出名为 Extreme Event Aware(η-learning)的机器学习方法,无需依赖历史极端事件数据即可生成合理的极端天气与最坏情景,并给出其规模、强度和持续时间。
Import AIAI 评分2828 Import AI 470:机器不应享有权利、SPADE 自动生成训练环境、Hawkeye 优化 GPU 内核
METR 研究显示 AI 对科学的加速并不均匀:2026 年网络安全漏洞报告速度较 2025 年大幅加快,数学领域贡献有限,AI 研究本身则未见可测量的加速。多校团队提出 SPADE 框架,让 LLM 交替生成可执行训练环境并求解,在 Qwen3-30B-A3B 上使游戏环境套件均分达 58.3,较基座提升 8.1。
Andrew Ng@AndrewYNgAI 评分5454引用Percy Liang@percyliang🚢 Marin 535B-A23B started training this week! As usual, the whole process is open. Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow. Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.
Chips and Cheese(RSS)AI 评分3434 Hot Chips 2026:High Bandwidth Flash(HBF)如何应用于机器学习负载
Hot Chips 2026 教程日上,Anurag Agarwal 与 Radhakrishna Giduthuri 探讨了 High Bandwidth Flash(HBF)在机器学习负载中的应用,由于尚无 HBF 产品,演讲基于模拟与预测。
Chips and Cheese(RSS)AI 评分3030 Hot Chips 2026:三星与 HBM Base Die 的机会
三星在 Hot Chips 2026 上介绍了 HBM Base Die 的三阶段演进方案,HBM4 与 HBM4E 的 base die 已改用 4nm 逻辑工艺以抑制功耗上升。
vLLM 官方博客(RSS)AI 评分4646 vLLM 如何用 Ray Direct Transport(RDT)实现大规模分片权重传输
vLLM 推出基于 Ray Direct Transport(RDT)的原生分片权重传输引擎,在 48 个 8xH100 节点(32 训练、16 推理)上以 BF16 完成 Kimi K2 的权重传输仅需 7.53 秒。
Noah Zweben@noahzwebenAI 评分2525SemiAnalysis 长文 RSS(RSS)AI 评分5757 SemiAnalysis 分析开源模型追赶闭源前沿的周期规律:每代耗时减半
SemiAnalysis 将 LLM 史分为早期扩展、推理和智能体三个时代,按时代分别用当时基准测算开源与闭源模型的综合能力分。
jietang@jietangAI 评分4646引用Liam Fedus@LiamFedusAn excellent history of scaling laws from @jietang. In 2020, we explored the limits of sparsity in Switch Transformers by routing each token to only 1 out of 2048 experts (in retrospect, a bold choice). The model had fewer than 3B activated parameters, but 1.6T total parameters (comparable to today's frontier models). The 1.6T model achieved better C4 perplexities than the T5 models using far less compute, set a new SOTA on TriviaQA, but was dumb as bricks on reasoning tasks like SuperGLUE. The lesson was that the optimal tokens-per-parameter ratio is highly task-dependent. Or as @NShazeer had already intuited: FLOPs were intelligence; parameters were knowledge!
vLLM 官方博客(RSS)AI 评分4242 SkyRL 推出 IsoExec:统一执行消除训练与推理引擎不匹配
SkyRL 提出 IsoExec,通过跨框架统一执行抽象消除 RL 训练与推理引擎之间的数值不匹配,包含执行契约与对齐的批不变内核两部分,已在 SkyRL 中结合 vLLM 和 Megatron 实现。
Microsoft Research 博客(RSS)AI 评分4141 微软 Skala 1.1 发布:训练数据增 2.5 倍,已接入 CP2K 并推进 Psi4、FHI-aims、ORCA、VASP 集成
微软研究院发布深度学习交换关联泛函 Skala 1.1,训练数据比首个公开版本多 2.5 倍,在 GMTKN55 的 55 个类别中 32 项排名第一,加权平均误差 2.8 kcal/mol,精度超过当前领先的全局(范围分离)杂化泛函而保持半局域泛函的计算成本。
NVIDIA Technical Blog(开发者技术博客 · RSS)AI 评分2525 生成式推荐系统如何重新定义大规模 RecSys
生成式推荐系统正把 RecSys 从传统嵌入向量相似度目标转向生成式目标,即根据用户历史序列预测大目录中的下一个动作或物品。NVIDIA 技术博客指出,RecSys 是消费互联网中最普遍却极难大规模训练与服务的机器学习问题之一,LLM 的出现推动了这一范式转变。
Jim Fan@DrJimFanAI 评分5252vLLM 官方博客(RSS)AI 评分4343 VeRL-Omni v0.2.0 发布:更快的扩散模型 RL 与稳定的全模态训练
VeRL-Omni v0.2.0 发布,通过 vLLM-Omni 的请求级批处理让 Qwen-Image FlowGRPO rollout 的 GPU 利用率从约 80% 升至约 100%,单次生成时间从 226s 降至 108s,减少 52%。
NVIDIA Technical Blog(开发者技术博客 · RSS)AI 评分3131 用 NVIDIA FLARE 构建联邦多模态 AI 工作流
NVIDIA 发布技术博客,介绍如何用 NVIDIA FLARE 构建联邦多模态 AI 工作流,在数据无法集中到一处的情况下跨数据本地站点协调视觉语言模型(VLM)训练。VLM 可支持视觉问答、图像描述和图文推理等任务,而联邦学习为这类分布式数据场景提供了训练协调方案。
jietang@jietangAI 评分5353MIT News(RSS)AI 评分6464 MIT CSAIL 研究:生成图像的归因随训练数据规模增大而衰减
MIT CSAIL 团队在 Nature Communications 发表论文,提出“归因衰减”现象:训练数据越大,单个样本对生成结果的影响越小,删除任一图像甚至某艺术家的全部图像后输出不变。
蚂蚁 inclusionAI:GitHub 新仓库AI 评分1919 蚂蚁 inclusionAI 发布 SingProbe 官方训练代码
蚂蚁 inclusionAI 在 GitHub 上线 SingProbe 官方训练代码仓库。正文仅说明这是 SingProbe 的官方训练代码,未披露模型规模、训练数据或评测结果等细节。
Ahead of AI(RSS)AI 评分5757 Sebastian Raschka 从零构建 AI 文本检测器教程
Sebastian Raschka 撰写教程,从零构建一个类似 Substack 内置 AI 检测功能的 AI 文本检测器,方法参考 Pangram 模型,通过微调 DistilBERT 分类器输出 0-100 的 AI 生成概率分。
Microsoft Research 博客(RSS)AI 评分3737 MindTopo 揭示多模态大模型的空间推理能力短板
微软研究院推出 MindTopo 基准,从连续性、分离、顺序、包围、绳结五类拓扑关系评估多模态大模型的推理与规划能力。测试显示,模型在静态图像识别上表现明显优于交互式规划任务,失败多发生在规划阶段而非感知阶段,且整体远低于人类水平。图像与视频生成仅在单帧关系可见时偶有帮助,跨多步动作时难以维持拓扑约束。
Chips and Cheese(RSS)AI 评分3333 Synopsys 谈芯片设计的物理学:从 3DIC 到热管理
Synopsys 的 Ravi Subramanian 在 DAC 2026 音频访谈中讨论了芯片设计中的物理学与 EDA 工具,重点谈及 3DIC 与热管理。他指出典型移动 SoC 约 2 到 25 亿门,而典型汽车 ECU 芯片约 70 亿门,功耗已直接影响到电动车续航。随着芯片变大,机械应力等原本的二三阶效应正变成一阶效应,签核需同时考虑机械与电学性能。
Microsoft Research 博客(RSS)AI 评分3333 Microsoft Research 推出 CARE-X:面向临床放射学的统一胸部 X 光 VLM
Microsoft Research 发布研究模型 CARE-X,一个统一胸部 X 光视觉语言模型,同时支持报告生成与结构化预测,并用强化学习(DAPO)在多任务设置下奖励临床正确性。
MIT News(RSS)AI 评分4444 MIT CSAIL 与清华提出 GeoPT:让 AI 模型学会物理,仿真提速 2 倍、数据省 60%
MIT CSAIL 与清华研究人员提出预训练方法 GeoPT,通过 130 万条"合成动力学"样本让仿真模型学习物理规律,达到峰值性能的速度比领先模型快 2 倍,所需数据最多减少 60%。在工业基准上,GeoPT 在速度、精度和效率上超越 SOTA 模型,模拟船体受风浪时用 60% 更少标注数据、达到峰值精度快 4 倍,并能在数秒内完成超 1 亿网格点的高保真仿真。