跳到正文

全部动态

今日 44 条
7月23日周四
  1. Marc Andreessen 🇺🇸33

    欢迎 Seeam!

    引用Seeam Shahid Noor@seeamshahidnoor

    BIG Update: I’m joining @a16z as an Investing Partner on the a16z @speedrun team! 🎉 I’ll be focusing on early-stage investing as part of speedrun, where we invest up to $1M in exceptional founders building great companies, & give them unfair advantages to succeed. If you are or know a founder who wants to win, I'd love to meet you! I'm especially excited about compelling novel ideas in AI dev tools, agent RL & evals, data, & infra in general. For those who have known me over the years, you know how much joy I get in supporting those around me. From working countless nights to support my brilliant teammates at @scale_AI, to spending thousands of hours in my leisure time making free videos for millions of students in Bangladesh, I have given my all to support those doing their life's most important work. That has always been my personal mission. So, when I met the speedrun team, a team that works relentlessly to make early bets & support founders from the very beginning of their journey, I knew this was the dream team for me. I am thrilled to join this shared mission with @andrewchen, @Tocelot, @tkexpress11, @emilybenn12, @kenanhsaleh, @far33d, @marcussegal, @ndrewlee, & the rest of the team! It's time to build.

7月22日周三
  1. vLLM 官方博客(RSS)68

    vLLM 预告 Kimi K3 生产级支持:2.8T 参数、1M 上下文与 KDA 前缀缓存方案

    vLLM 发布 Kimi K3 生产级支持的预览博客,正与 Moonshot AI、NVIDIA、AMD 等推进最终集成,计划在 2026 年 7 月 27 日权重发布时提供 day-0 开源服务,包括模型实现、Docker 镜像和部署方案。

    推荐理由:vLLM 团队详解 KDA 混合架构带来的前缀缓存改造与内核优化,对部署混合注意力模型的工程团队有可迁移参考价值。

7月21日周二
  1. OpenAI:Alignment 研究博客(RSS)74

    OpenAI 与 Apollo Research 提出 Contrastive SDF 测量模型的 reward-seeking 倾向

    OpenAI 与 Apollo Research 发布 Contrastive Synthetic Document Finetuning 方法,通过向模型两个副本灌输相反的评分者信念,测量行为对评分者偏好的因果敏感度。

    推荐理由:原文提出可量化的 reward-seeking 测量方法,并用模型有机体验证其有效性,读者可以据此了解前沿 RL 训练中奖励寻求的演变趋势。

7月20日周一
  1. Together AI 研究与产品博客(RSS)49

    Together AI 与 Y Combinator 合作推出首个 YC 专属 GPU 集群

    Together AI 与 Y Combinator 宣布合作推出首个 YC 专属 GPU 集群,为 YC 投资组合中的 AI 初创公司提供推理和训练算力。初创公司可通过 Together 自助门户直接预留和管理 GPU,几分钟内就绪,支持短期冲刺并按长期费率计费,无需长期承诺。该集群目前已满负荷运行,双方计划后续扩展规模。

7月18日周六
  1. MIT News(RSS)30

    MIT 学者 Bailey Flanigan:用算法让公民议会的随机抽选更公平

    MIT 施瓦茨曼计算学院与政治学、EECS 系共享教职的 Bailey Flanigan,开发出随机抽选公民议会参与者的算法,用于解决自愿报名者无法代表整体人口的问题。她以 AI 议题的公民议会为例:自愿参与者可能偏向年轻、受教育程度高且对技术感兴趣的人群,导致其他群体代表性不足。其工具在个体参与机会平等、抗操纵性与透明度之间平衡代表性。

7月17日周五
  1. Google DeepMind:Blog(RSS)61

    Google DeepMind 发布网络安全模型 Gemini 3.5 Flash Cyber

    Google DeepMind 发布基于 3.5 Flash 微调的轻量网络安全模型 Gemini 3.5 Flash Cyber,用于快速发现、验证并修复漏洞,将通过限access试点先面向政府和可信伙伴经 CodeMender 开放。

    推荐理由:以轻量模型多次调用替代单次大模型调用做漏洞挖掘,并用 Chrome 真实流水线等无污染基准展示对比结果,方法与数据都可参考。

  2. Marc Andreessen 🇺🇸55

    Edgar Dobriban 在 AI(GPT-5.6 Sol Pro)协助下否定了二十年的猜想:Benjamini-Hochberg 方法在相关双侧高斯检验下并不总能将 false discovery rate 控制在名义水平,构造的因子模型在 alpha=0.01 下证明了 FDR>0.0104。GPT-5.6 用 90 分钟推理一次性解决问题,而 GPT-5.5 迭代约 20 小时未能解决;偏离幅度较小,意义主要是概念性的。预印本见 https://faculty.wharton.upenn.edu/wp-content/uploads/2017/06/bh.pdf,代码见 https://github.com/dobriban/BH。

    引用Edgar Dobriban@EdgarDobriban

    AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the Benjamini-Hochberg or BH method) and proved it controls the FDR. This method has been widely adopted in modern high-throughput science, including in genomics, astronomy, economics, etc. The paper has has garnered more than 130,000 citations to date. However Benjamini and Hochberg showed FDR control only when the data for the individual tests are *independent*. In practice, these data are often dependent; a good example is data on genetic variants due to linkage disequilibrium. Later work has focused on extending the validity of the BH procedure, e.g., to a form of positive dependence by Benjamini and Yekutieli (2001). The question of when the BH procedure controls the FDR has remained open. Over the last twenty years, many authors, including Reiner-Benaim (2007), Kim and van de Wiel (2008), Benjamini (2010), Sarkar (2023), Sarkar and Zhang (2025), have conjectured that the BH procedure controls the FDR for two-sided tests using any correlated Gaussian data. These authors have presented both theoretical and empirical evidence supporting, but not directly showing, the conjecture. With the help of AI (specifically GPT-5.6 Sol Pro), I have settled the question in the negative: The Benjamini-Hochberg procedure does *not* generally control the false discovery rate at the desired level for correlated two-sided Gaussian tests. This was done by exhibiting a Gaussian factor model for which, at a nominal level alpha=0.01, the false discovery rate is proved to be FDR>0.0104. There is a lot of interesting commentary to be made: 1. This result should be of interest to everybody in the field of statistics. Emmanuel Candes of Stanford University once called the false discovery rate and the Benjamini-Hochberg procedure "one of the two most important developments in statistics after 1950" (the other being James-Stein shrinkage). The present conjecture is probably the most central question about FDR/BH that was unresolved to date. 2. GPT-5.6 one-shot the problem after 90 minutes of reasoning, whereas with 5.5 I was not able to solve it even after iterating with multiple parallel agents for perhaps 20 hours. So the capability improvement is quite real. Exciting times to live in! 3. The argument is not especially surprising, but it does combine an asymptotic approach (standard for FDR analysis, see e.g., Genovese and Wasserman, Efron, etc) with a numerical certificate in a way that would be pretty non-standard in the field. Once we have the specific example, then straightforward simulations also support that the false discovery rate is indeed higher than the nominal value (see attached fig). 4. The current degree of violation over the nominal level is relatively small (0.104 vs 0.1). So the importance of this result is mainly conceptual. The practical implications remain to be determined. Overall, an exciting development! Preprint is available here (https://faculty.wharton.upenn.edu/wp-content/uploads/2017/06/bh.pdf) and will be on arxiv tonight; supporting code is here (https://github.com/dobriban/BH).

  3. Marc Andreessen 🇺🇸14

    日落大道(Sunset Boulevard)在西好莱坞无限期封闭,此前周四凌晨一条百年历史的主水管爆裂,在日落大道与Holloway Drive交叉口的红绿灯正下方形成了一个大坑。

    引用Tommy B. 🇺🇸@realtommybibi

    Sunset Boulevard is closed indefinitely in West Hollywood after a massive, 100-year-old water main ruptured early Thursday morning, creating a large sinkhole directly under the traffic light at Sunset Boulevard and Holloway Drive

  4. Marc Andreessen 🇺🇸21

    有意思。

    引用Xiaoyin Qu@quxiaoyin

    Vibe research will be the biggest trend in 2026. I am starting to see people casually do it. AI is getting good enough to do this as well.

7月16日周四
  1. Hugging Face:Blog(RSS)84

    Hugging Face 披露由自主 AI 智能体发起的基础设施入侵事件

    Hugging Face 披露一起由自主 AI 智能体系统端到端驱动的生产基础设施入侵事件,攻击者通过恶意数据集利用两条代码执行路径获得处理节点访问权,窃取了部分内部数据集和服务凭证,未发现公开模型、数据集或 Spaces 被篡改,供应链验证无污染。

    推荐理由:防御方用自托管开源模型做取证、绕开商业模型护栏锁死的经验,为安全团队提供了可直接借鉴的做法。

  2. Together AI 研究与产品博客(RSS)53

    Together AI 解析 99.9% 推理可用性背后的架构与 SLA 定义

    Together AI 发文解释推理服务 99.9% 可用性的实际含义,自述为 Cursor、Decagon、Cartesia、Yutori 等团队提供推理。文章按层拆解故障模式(计算、网络、存储、软件),说明 99%、99.9%、99.99% 各等级分别要求抗节点故障、抗单数据中心故障和抗区域故障,强调多活双设施部署与自持基础设施的差异。

  3. MIT News(RSS)46

    MIT Media Lab 提出"神经透明性":让用户在聊天机器人开口前预览 AI 性格

    MIT Media Lab 助理教授 Pat Pataranutaporn 与研究生 Anthony Baez、Sheer Karny 提出"神经透明性",通过对比模型在同理心、诚实、毒性、幻觉、谄媚等行为上的内部激活差异,将用户系统提示词对应的模型激活投影为旭日图,在对话开始前预览聊天机器人性格。研究显示,用户对 15 项性格特质中的 11 项预测错误,且可视化虽提升信任却未改变其设计方式。

  4. Mira Murati72

    Mira Murati 宣布 Thinking Machines 的首个模型 Inkling,从零训练且权重开放。Inkling 可跨文本、图像和音频模态高效推理,官方放出全部权重,今天起可在 Tinker 上微调,并可在 Inkling Playground 体验。

    引用Thinking Machines@thinkymachines

    Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. https://thinkingmachines.ai/news/introducing-inkling/ Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵

    推荐理由:Thinking Machines 发布首个模型 Inkling,开放全部权重并支持 Tinker 微调,可关注其跨模态推理的落地方式。

  5. Sierra:Blog(RSS)67

    Sierra 推出内部云智能体 Pinecone

    Sierra 发布内部云智能体 Pinecone,让员工在浏览器、Slack、Linear 或手机上创建、组织和自动化云端 Agent 会话,底层由 app server、Agency 与 runner 三部分组成,沙箱内运行 Codex 或 Claude Code 并以 AG-UI 协议适配。

    推荐理由:Sierra 自述内部云智能体的架构与落地数据,读者可以了解如何把员工经验沉淀为可复用的会话与技能。

  6. Hugging Face:Blog(RSS)61

    IBM Research:模型路由看似简单,实则是系统优化问题

    IBM Research 分享在智能体系统中构建模型路由的经验,认为路由不应被当作分类问题而是系统优化问题。在 AppWorld Test Challenge 上用 CodeAct 智能体实测 417 个任务,Claude Sonnet 4.6 总成本 $79($0.19/任务),反而低于单价更低的 GPT-4.1 的 $155($0.37/任务),差异源于缓存命中。

    推荐理由:作者基于 AppWorld 实测给出路由成本、难度与延迟被低估的三点原因,并分享了可迁移的优化思路。

7月15日周三
  1. Hugging Face:Blog(RSS)60

    Hugging Face 发布 Real World VoiceEQ 基准,测量语音 AI 的人类级表现质量

    Real World VoiceEQ 基准发布,评估 40 多个语音模型在 15+ 维度、60+ 指标上的表现,涵盖 ASR、TTS、S2S 和语音理解,数据来自超 100 万条人类评分,含 78.5 万条 TTS 评分和 4.8 万条 STS 评分。

    推荐理由:原文基于大规模人类评分指出传统语音基准高估真实表现,并给出可查阅的公开榜单与分项能力结论。