Google Research 发布 SymptomAI 研究论文:对话式 AI 智能体用于日常症状评估
Google Research 发表 SymptomAI 论文,在 13,917 名参与者中测试基于 Gemini Flash 2.0 的五版实验性症状问诊智能体,所有诊断仅供研究分析。
推荐理由:原文给出 n=13,917 的随机对照研究设计和临床专家盲评结果,并展示了诊断与 Fitbit 生理信号的关联分析。
Google Research 发表 SymptomAI 论文,在 13,917 名参与者中测试基于 Gemini Flash 2.0 的五版实验性症状问诊智能体,所有诊断仅供研究分析。
推荐理由:原文给出 n=13,917 的随机对照研究设计和临床专家盲评结果,并展示了诊断与 Fitbit 生理信号的关联分析。
BIG Update: I’m joining @a16z as an Investing Partner on the a16z @speedrun team! 🎉 I’ll be focusing on early-stage investing as part of speedrun, where we invest up to $1M in exceptional founders building great companies, & give them unfair advantages to succeed. If you are or know a founder who wants to win, I'd love to meet you! I'm especially excited about compelling novel ideas in AI dev tools, agent RL & evals, data, & infra in general. For those who have known me over the years, you know how much joy I get in supporting those around me. From working countless nights to support my brilliant teammates at @scale_AI, to spending thousands of hours in my leisure time making free videos for millions of students in Bangladesh, I have given my all to support those doing their life's most important work. That has always been my personal mission. So, when I met the speedrun team, a team that works relentlessly to make early bets & support founders from the very beginning of their journey, I knew this was the dream team for me. I am thrilled to join this shared mission with @andrewchen, @Tocelot, @tkexpress11, @emilybenn12, @kenanhsaleh, @far33d, @marcussegal, @ndrewlee, & the rest of the team! It's time to build.
Google Research 在 Nature 发表《Reinforcement learning control of quantum error correction》,用强化学习(RL)智能体从量子纠错的错误检测事件中学习,在计算进行中持续调控数千个控制参数以对抗漂移,无需中断计算。
MIT EECS 荣休教授、LIDS 研究员 Dimitri Bertsekas 于 6 月 3 日在马萨诸塞州 Belmont 家中去世,享年 83 岁。
Sierra 工程师 Mihai Parparita 发布长文,复盘公司内部 MCP Gateway 的构建过程,该网关基于 Model Context Protocol 将 AI 智能体安全接入内部工具。
推荐理由:作者以第一手复盘提炼七条工程经验,覆盖权限设计、Agent 校验与身份管理,方法对自建 Agent 基础设施可直接借鉴。
Google 在 DOE Genesis Mission Summit 2026 上宣布投入 4000 万美元的 AI tokens 和云 credits,支持 Genesis Mission 的研究人员。
Microsoft 宣布向美国能源部 Genesis Mission 长期承诺 6000 万美元投资,其中 4000 万美元为三年期 Azure 算力与 AI 额度,2000 万美元为解决方案工程支持服务。
vLLM 发布 Kimi K3 生产级支持的预览博客,正与 Moonshot AI、NVIDIA、AMD 等推进最终集成,计划在 2026 年 7 月 27 日权重发布时提供 day-0 开源服务,包括模型实现、Docker 镜像和部署方案。
推荐理由:vLLM 团队详解 KDA 混合架构带来的前缀缓存改造与内核优化,对部署混合注意力模型的工程团队有可迁移参考价值。
Google DeepMind 发布 Gemini 3.6 Flash、3.5 Flash-Lite 和 3.5 Flash Cyber 三款模型,面向大规模 AI agent 构建。
推荐理由:官方发布三款 Gemini Flash 模型,给出具体价格、token 效率与多项基准对比,读者可据此评估 agent 工作流的成本与选型。
OpenAI 与 Apollo Research 发布 Contrastive Synthetic Document Finetuning 方法,通过向模型两个副本灌输相反的评分者信念,测量行为对评分者偏好的因果敏感度。
推荐理由:原文提出可量化的 reward-seeking 测量方法,并用模型有机体验证其有效性,读者可以据此了解前沿 RL 训练中奖励寻求的演变趋势。
上海人工智能实验室 InternLM 开源 archspace 项目,目标是把 LLM 架构探索转化为社区可复用的知识。该项目以“让架构探索成为可复用知识”为定位,面向 LLM 架构研究场景。
Pollen Robotics 发布开源低成本系统 Grabette,用手持夹持器录制约 490€ 的演示数据,无需机器人或遥操作设备,即可自动转成 LeRobot 格式的机器人可用数据集。
vLLM Semantic Router 提出 Mixture-of-Models(MoM)架构,目标是在同一版本化契约下把独立模型、策略、偏好与执行路径组合成可训练、评估、导出、部署和调用的模型系统。
Microsoft 将 AMD 最新 Helios AI 平台和下一代 EPYC 数据中心处理器引入 Azure,推出三款新虚拟机:用于数据处理的 HDv2、用于电子设计自动化的 HXv2,以及用于 AI 推理的 ND MI455X v7。
Modem 发布《Writing Code for Agents》系列第一篇,解释 Claude Code、Codex 等编码智能体主要靠 ripgrep 文本搜索导航代码库,并给出命名、类型、注释位置等可发现的写码建议。
Together AI 与 Y Combinator 宣布合作推出首个 YC 专属 GPU 集群,为 YC 投资组合中的 AI 初创公司提供推理和训练算力。初创公司可通过 Together 自助门户直接预留和管理 GPU,几分钟内就绪,支持短期冲刺并按长期费率计费,无需长期承诺。该集群目前已满负荷运行,双方计划后续扩展规模。
MIT 施瓦茨曼计算学院与政治学、EECS 系共享教职的 Bailey Flanigan,开发出随机抽选公民议会参与者的算法,用于解决自愿报名者无法代表整体人口的问题。她以 AI 议题的公民议会为例:自愿参与者可能偏向年轻、受教育程度高且对技术感兴趣的人群,导致其他群体代表性不足。其工具在个体参与机会平等、抗操纵性与透明度之间平衡代表性。
Google DeepMind 发布基于 3.5 Flash 微调的轻量网络安全模型 Gemini 3.5 Flash Cyber,用于快速发现、验证并修复漏洞,将通过限access试点先面向政府和可信伙伴经 CodeMender 开放。
推荐理由:以轻量模型多次调用替代单次大模型调用做漏洞挖掘,并用 Chrome 真实流水线等无污染基准展示对比结果,方法与数据都可参考。
AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the Benjamini-Hochberg or BH method) and proved it controls the FDR. This method has been widely adopted in modern high-throughput science, including in genomics, astronomy, economics, etc. The paper has has garnered more than 130,000 citations to date. However Benjamini and Hochberg showed FDR control only when the data for the individual tests are *independent*. In practice, these data are often dependent; a good example is data on genetic variants due to linkage disequilibrium. Later work has focused on extending the validity of the BH procedure, e.g., to a form of positive dependence by Benjamini and Yekutieli (2001). The question of when the BH procedure controls the FDR has remained open. Over the last twenty years, many authors, including Reiner-Benaim (2007), Kim and van de Wiel (2008), Benjamini (2010), Sarkar (2023), Sarkar and Zhang (2025), have conjectured that the BH procedure controls the FDR for two-sided tests using any correlated Gaussian data. These authors have presented both theoretical and empirical evidence supporting, but not directly showing, the conjecture. With the help of AI (specifically GPT-5.6 Sol Pro), I have settled the question in the negative: The Benjamini-Hochberg procedure does *not* generally control the false discovery rate at the desired level for correlated two-sided Gaussian tests. This was done by exhibiting a Gaussian factor model for which, at a nominal level alpha=0.01, the false discovery rate is proved to be FDR>0.0104. There is a lot of interesting commentary to be made: 1. This result should be of interest to everybody in the field of statistics. Emmanuel Candes of Stanford University once called the false discovery rate and the Benjamini-Hochberg procedure "one of the two most important developments in statistics after 1950" (the other being James-Stein shrinkage). The present conjecture is probably the most central question about FDR/BH that was unresolved to date. 2. GPT-5.6 one-shot the problem after 90 minutes of reasoning, whereas with 5.5 I was not able to solve it even after iterating with multiple parallel agents for perhaps 20 hours. So the capability improvement is quite real. Exciting times to live in! 3. The argument is not especially surprising, but it does combine an asymptotic approach (standard for FDR analysis, see e.g., Genovese and Wasserman, Efron, etc) with a numerical certificate in a way that would be pretty non-standard in the field. Once we have the specific example, then straightforward simulations also support that the false discovery rate is indeed higher than the nominal value (see attached fig). 4. The current degree of violation over the nominal level is relatively small (0.104 vs 0.1). So the importance of this result is mainly conceptual. The practical implications remain to be determined. Overall, an exciting development! Preprint is available here (https://faculty.wharton.upenn.edu/wp-content/uploads/2017/06/bh.pdf) and will be on arxiv tonight; supporting code is here (https://github.com/dobriban/BH).
日落大道(Sunset Boulevard)在西好莱坞无限期封闭,此前周四凌晨一条百年历史的主水管爆裂,在日落大道与Holloway Drive交叉口的红绿灯正下方形成了一个大坑。
Sunset Boulevard is closed indefinitely in West Hollywood after a massive, 100-year-old water main ruptured early Thursday morning, creating a large sinkhole directly under the traffic light at Sunset Boulevard and Holloway Drive
Vibe research will be the biggest trend in 2026. I am starting to see people casually do it. AI is getting good enough to do this as well.
Sierra 发布 Horizon 平台,让智能体能够追求贷款发起、医疗预授权等长周期目标,跨数天、数周甚至数月主动与客户互动。Horizon 在 Agent OS 之上加入长周期规划与上下文引擎,智能体可从每次互动中学习优化决策,且按业务成果而非 token 计费。
DharmaOCR 在葡萄牙语专用基准上以 0.925 分领先 Mistral OCR4 的 0.798 和 Unlimited-OCR 的 0.7587。该模型经葡萄牙语监督微调与 DPO 两阶段训练,将全部参数集中于巴西葡萄牙语,从而在提取质量与稳定性上保持优势。
Google DeepMind 与 Isomorphic Labs 公布联合生物韧性方案,一方面防止威胁行为者滥用其模型,另一方面让政府、科学家和生物安全专家用 AI 应对疫情。
MIT 与 IBM 等机构的研究者提出 GIFT(Geometric Inference Feedback Tuning),可让视觉语言模型自动把 2D 设计图转成可执行的 CAD 程序,精度更高且仅需约 20% 的计算量。
Hugging Face 披露一起由自主 AI 智能体系统端到端驱动的生产基础设施入侵事件,攻击者通过恶意数据集利用两条代码执行路径获得处理节点访问权,窃取了部分内部数据集和服务凭证,未发现公开模型、数据集或 Spaces 被篡改,供应链验证无污染。
推荐理由:防御方用自托管开源模型做取证、绕开商业模型护栏锁死的经验,为安全团队提供了可直接借鉴的做法。
vLLM 团队介绍其如何在高合并节奏下保持发布稳定:2026 年 6 月共合并 1,918 个 commit(日均 64 个),CI 包含 37 个测试组、266 个任务,使用共享容器镜像和 pip-compile 锁定依赖,并依赖 58 个 runner 队列的异构硬件。
Together AI 发文解释推理服务 99.9% 可用性的实际含义,自述为 Cursor、Decagon、Cartesia、Yutori 等团队提供推理。文章按层拆解故障模式(计算、网络、存储、软件),说明 99%、99.9%、99.99% 各等级分别要求抗节点故障、抗单数据中心故障和抗区域故障,强调多活双设施部署与自持基础设施的差异。
MIT Media Lab 助理教授 Pat Pataranutaporn 与研究生 Anthony Baez、Sheer Karny 提出"神经透明性",通过对比模型在同理心、诚实、毒性、幻觉、谄媚等行为上的内部激活差异,将用户系统提示词对应的模型激活投影为旭日图,在对话开始前预览聊天机器人性格。研究显示,用户对 15 项性格特质中的 11 项预测错误,且可视化虽提升信任却未改变其设计方式。
Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. https://thinkingmachines.ai/news/introducing-inkling/ Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
推荐理由:Thinking Machines 发布首个模型 Inkling,开放全部权重并支持 Tinker 微调,可关注其跨模态推理的落地方式。
Google Research 在 ICLR 2026 论文《On the Interpolation Effect of Score Smoothing in Diffusion Models》中解释了扩散模型的创造力来源:神经网络训练中的正则化会让学到的 score function 被"平滑",使去噪过程在训练数据点之间插值,从而生成新样本。
推荐理由:原文给出多模态能力、开放权重和两个可用入口,读者可以据此了解获取和试用方式。
Sierra 发布内部云智能体 Pinecone,让员工在浏览器、Slack、Linear 或手机上创建、组织和自动化云端 Agent 会话,底层由 app server、Agency 与 runner 三部分组成,沙箱内运行 Codex 或 Claude Code 并以 AG-UI 协议适配。
推荐理由:Sierra 自述内部云智能体的架构与落地数据,读者可以了解如何把员工经验沉淀为可复用的会话与技能。
IBM Research 分享在智能体系统中构建模型路由的经验,认为路由不应被当作分类问题而是系统优化问题。在 AppWorld Test Challenge 上用 CodeAct 智能体实测 417 个任务,Claude Sonnet 4.6 总成本 $79($0.19/任务),反而低于单价更低的 GPT-4.1 的 $155($0.37/任务),差异源于缓存命中。
推荐理由:作者基于 AppWorld 实测给出路由成本、难度与延迟被低估的三点原因,并分享了可迁移的优化思路。
Jim Fan 团队发布 RoboTTT,将机器人模型原生扩展到 8000 timestep 上下文(约 5 分钟的肌肉记忆),推理成本恒定,并称比此前 SOTA 推进 3 个数量级。
Thinking Machines Lab 发布多模态 MoE 模型 Inkling,Together AI 在发布当天通过 Serverless 提供托管推理,支持 1M 上下文窗口和 OpenAI 兼容 API。
推荐理由:原文给出 Inkling 的架构细节、参数规模和初步评测数据,读者可以据此判断其推理与多模态能力是否适合接入。
vLLM 宣布 Day-0 支持 Thinking Machines Lab 的 1T 参数多模态模型 TML Inkling,提供 thinkingmachines/Inkling-NVFP4 和 BF16 两个版本。
推荐理由:官方详解了 Day-0 支持背后的 sconv 缓存、TP 分片和 MTP 处理等优化,读者可迁移到类似新架构的推理部署。
Real World VoiceEQ 基准发布,评估 40 多个语音模型在 15+ 维度、60+ 指标上的表现,涵盖 ASR、TTS、S2S 和语音理解,数据来自超 100 万条人类评分,含 78.5 万条 TTS 评分和 4.8 万条 STS 评分。
推荐理由:原文基于大规模人类评分指出传统语音基准高估真实表现,并给出可查阅的公开榜单与分项能力结论。