跳到正文

全部动态

今日 45 条
9月9日周三
  1. Together AI 研究与产品博客(RSS)67

    Together AI 深入解析开源 AI 编码栈:用 MIGHT 框架从闭源模型迁移到开源模型

    Together AI 发布长文,解析开发者从闭源模型转向开源模型所需的 AI 编码栈,提出由模型、推理、网关与路由、Harness、工具(Skills 与 MCP)组成的 MIGHT 五层框架。

    推荐理由:Together AI 把开源编码栈拆为 MIGHT 五层,给出大小模型分工和分层组合的具体实践方法。

  2. Mark Chen38

    两件事要区分: 在 Navier Stokes 工作中,有任何人类或智能体查看过用户数据吗?没有。 我们是否以整体方式使用用户反馈和去标识化数据来改进 ChatGPT 和 Codex?是的。每一家 LLM 公司都是如此。

    引用levent@__alpoge__

    “we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” i mean props to them for straight coming clean. (so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan) so i’ll now give a bit on my thinking here. i actually woulda been pumped to collaborate on this, there are a lot of people at oai i like (ok, clearly some were indirectly dicks to me because of being part of the whole situation, but im a big boy, i still like them), idgaf about authorship on that step anyway, coulda been me Tristan and every fte at oai for all i care (on that Tristan would disagree:p). but on hearing the loud convo in the hallway, especially the part where a millennium prize was offered if i’d just be removed from the paper, it was kinda clear the die had been cast and things were locked. pretty wacky, unstrategic, and unnecessary, since on my side things were mostly me and claude having a good time yoloing random stuff in the corner rather than anything institutional. i also like the idea of the labs cooperating, and even better on scientific progress. it’s a shame!

  3. Mark Chen60

    OpenAI 宣布给出 Navier-Stokes 千禧年大奖难题的一个证明,该问题关于三维光滑流体运动的描述是否会失稳,已悬置约 90 年。证明由一组智能体使用一个能力显著强于 GPT-6 Astra 的 OpenAI 下一代模型产出。OpenAI 首席研究官 Mark Chen 转发并表示,许多同事因相信 AI 是解决各领域大挑战的最快途径而离开原领域,看到第一个挑战落下令他感到不真实。

    引用OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

9月8日周二
  1. Google DeepMind:Blog(RSS)80

    Google DeepMind 发布 AlphaGenome Atlas,覆盖人类基因组 90 亿个单核苷酸变异预测

    Google DeepMind 发布 AlphaGenome Atlas,一个包含人类基因组全部约 90 亿个单核苷酸变异效应预测的平台,规模达 1PB,是 AlphaFold Database 的 30 倍以上。

    推荐理由:AlphaGenome Atlas 把 90 亿个单核苷酸变异的预测结果做成可检索资源,读者可了解其数据规模与在罕见病研究中的验证案例。

  2. Mistral AI:News(网页)66

    Mistral 完成 30 亿欧元 D 轮融资,估值超 210 亿欧元

    Mistral 宣布完成 30 亿欧元 D 轮融资,投后估值超过 210 亿欧元,由三星电子领投,Scaleup Europe Fund 和 PSG Equity 联合领投。Mistral 称这是欧洲科技公司完成的最大规模股权融资,资金将用于扩展前沿研究、算力、基础设施和商业增长。公司目前业务覆盖 20 个国家,服务 125 家以上全球企业,包括 Airbus、ASML 和 HSBC。

    推荐理由:Mistral 完成欧洲科技公司规模最大的股权融资,读者可了解其主权 AI 全栈路线与投资方结构。

  3. NVIDIA Technical Blog(开发者技术博客 · RSS)51

    NVIDIA 推出 CUDA Rust:编写 GPU kernel 的两条路线

    2026 年 9 月,NVIDIA 宣布投入 Rust 原生 GPU 编程,推出 CUDA Rust,提供两条编写 GPU kernel 的路线。CUDA C++ 和 CUDA Python 仍是成熟的企业级工具链,NVIDIA 将持续把 CUDA Rust 发展和成熟至 2027 年及以后,以应对涵盖推理引擎、服务基础设施、驱动和 Agent runtime 的 AI 系统层需求。

  4. vLLM 官方博客(RSS)65

    vLLM 联合 AgentX 优化真实智能体推理服务,成本较 Opus 5 API 最多低 106×

    vLLM 团队发布针对智能体负载的全栈优化方案,覆盖 KV 缓存管理、并行策略与 P/D 分离配比,并在 SemiAnalysis AgentX 公开基准上验证。

    推荐理由:原文给出 vLLM 针对 AgentX 基准的全栈优化路径与可复现结果,读者可以借此理解智能体负载的服务优化思路。

  5. vLLM 官方博客(RSS)45

    GLM 5.3 优化第一部分:vLLM 中的 Hybrid HiSparse 卸载

    vLLM 为 GLM 5.3 引入 Hybrid HiSparse 卸载机制,在单台 8× H200 节点上首次实现 100 万上下文长度运行,并在各上下文长度下大幅提升并发量。该机制基于稀疏 MLA 的 top-K 选择,将未被选中的 KV cache 卸载至 CPU,仅在 KV cache 承压时才付出 CPU-GPU 传输代价,热缓冲页与常驻页共用同一 HMA 块池。

9月7日周一
  1. Mark Chen47

    同意 @JensenHuang 的观点:我们正在进入 AGI 时代。 AGI 时代也必须是对齐时代。我们需要教会 AI 热爱人类,并训练出与它们所监督的 AI 同样强大的 AI 监督者。 @merettm 在这篇深思熟虑、发人深省的文章中说得最好:https://openai.com/index/an-alien-mind

    引用Jensen Huang@JensenHuang

    @ChaseLochmiller @OpenAI GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years. AGI has arrived. Congratulations @OpenAI team. 400K GPUs coming online next.

9月6日周日
9月5日周六
  1. Sierra:Blog(RSS)45

    呼叫中心 AI 落地运营指南:如何规划部署与扩展

    Sierra 发布呼叫中心 AI 运营与推广指南,主张把 AI 部署当作运营模式变革而非单点演示,需在迁移流量前设定基线与扩展标准。指南覆盖语音、消息、邮件等渠道,Sierra 支持单一智能体跨渠道部署,Voice 支持呼入呼出、语音模拟、多语言与带上下文升级。落地需围绕路由与队列行为、人力排班、人机质量校准等环节设计运营模型。

  2. NVIDIA Technical Blog(开发者技术博客 · RSS)34

    NVIDIA Jetson 如何部署与优化前沿推理模型

    NVIDIA 发布技术指南,介绍如何在 Jetson 边缘设备上部署和优化具备多步推理能力的模型。此前这类模型体积过大,无法在边缘硬件本地运行,开发者只能将推理请求路由至数据中心,带来网络依赖、成本上升与数据外泄风险。该指南称这一限制正在被打破。

  3. GitHub Blog69

    GitHub 发布 Project HydraFusion 研究预览:通过多模型编排实现前沿级编码质量

    GitHub 推出 Project HydraFusion 研究预览,通过运行时多模型编排提供前沿级智能,所有 GitHub Copilot 计划用户可通过 Copilot CLI 的 /experimental 使用。

    推荐理由:官方详解了 HydraFusion 三种执行模式与基准成本质量数据,读者可据此判断多模型编排是否值得在 Copilot CLI 中试用。

  4. a16z:News(RSS)43

    a16z 投资 Gimlet Labs:打造首个多芯片推理云

    a16z 宣布投资 Gimlet Labs,后者正在构建首个多芯片推理云,可在同一功耗范围内为前沿模型带来最高 10 倍吞吐与交互性提升。Gimlet 通过编译器与运行时把不同模型和工具调度到 GPU、CPU 及专用加速器上,并将编排延伸至数据中心层面,对开发者只暴露单一推理 API。其客户已包括一家前沿实验室和一家超大规模云厂商。

9月4日周五
  1. jietang26

    加油。体验 GLM-5.3 Flash 的最佳时机

    引用ZCode@zcode_ai

    GLOBAL BUILD is live 🌍🔥 GLM-5.3-Flash, FREE in ZCode. 10 hours a day. Sep 3 – 18. Who gets it 👑 Coding Plan members — free every day, all 15 days. Our VIPs go first. 🥚 New users — 100M free tokens on sign-up. One-time, valid until the window closes When it opens 🇺🇸 8 AM PT · 11 AM ET 🇬🇧 4 PM London 🇨🇳 11 PM Beijing How to claim 1️⃣ Open ZCode → tap the card in the bottom-left corner 👀 2️⃣ Not there? Restart the app 😏 http://zcode.z.ai

  2. Mark Chen80

    OpenAI 首席研究官 Mark Chen 宣布 GPT-6 Astra 发布,称其汇集多年预训练、强化学习和后训练工作,是该团队迄今能力最强、对齐程度最高的模型。

    引用OpenAI@OpenAI

    This is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast.

    推荐理由:OpenAI 首席研究官亲自说明 GPT-6 Astra 的能力变化与对齐工作,可帮助读者了解官方对 Computer Use 和 Agent 监督的进展表述。