Tomer Tunguz 解析 GPU 租金翻倍而 AI 成本仍下降的原因
GPU 租赁价六个月内从 $4.40 涨到 $8.08 每卡时,但 AI 价格仍在下降。文章归因于数据中心建设推高电力、材料和信贷成本,需求激增放大竞价;同时模型效率快速提升,Claude Opus 5.5 运行成本降 40%,同一基准的通过成本从 $0.55 降至 $0.0015。Microsoft 称每 GPU 生成的 token 年增 90%,效率与成本两股力量大致相当。
GPU 租赁价六个月内从 $4.40 涨到 $8.08 每卡时,但 AI 价格仍在下降。文章归因于数据中心建设推高电力、材料和信贷成本,需求激增放大竞价;同时模型效率快速提升,Claude Opus 5.5 运行成本降 40%,同一基准的通过成本从 $0.55 降至 $0.0015。Microsoft 称每 GPU 生成的 token 年增 90%,效率与成本两股力量大致相当。
Google 团队在 MaxText(JAX/XLA)中从零复现 Ai2 的 Olmo 3 7B,在 Google Cloud TPU 上完成 stage-1 预训练(约5.93T token。
HeadGuard 是一种可组合的注意力头保护方法,通过离线选取约 1/8 的图像敏感与输出敏感 KV 头并保持其图像 key(可选 value)为 BF16 精度,来缓解低比特 KV-Cache 量化对 VLM 精度的损害。
DeepSeek 在 GitHub 上线新仓库 DeepEP-Ascend,这是一个面向华为昇腾 NPU 的高性能通信库,用于机器学习训练与推理。该库延续 DeepEP 的通信优化方向,将支持范围扩展至昇腾 NPU 平台。
Today we’re announcing OpenClaw Enterprise In collaboration with @RedHat , @nvidia and @OpenAI the OpenClaw Foundation is open sourcing a powerful enterprise control plane for persistent agents OpenClaw Enterprise is built to run on your own infrastructure and will always be free for an organization to use https://openclaw.ai/blog/openclaw-enterprise
NVIDIA 开源了基于 TensorRT 的 C++ AI 模型参考实现集合 TensorRT Model Connect,目标是让 NVIDIA 推理栈的性能更易获取。项目围绕编码智能体设计,采用并行工作、模型族隔离、可逆变更与 GPU 验证等实践。
Baseten joined the @OpenAI B2B marketplace today as one of the first open-model inference providers. The best AI companies are already running a mix of closed and open models at huge scale. We're excited to give OpenAI enterprise customers the ability to use open models powered by Baseten natively within Codex and through the Responses API.
We just raised an $8M seed round to kill AWS, GCP, and Azure. Introducing http://instacloud.com, the agent-native serverless cloud. Your team is shipping code like never before. But you're getting caught up in manual, tedious DevOps work trying to deploy it. InstaCloud provides the serverless compute that lets your services autoscale, with all the infrastructure managed for you. Agents branch into complete replica environments when working, keeping prod safe and iteration speed high. And of course, it all works seamlessly with agents through MCP/CLI. Get off the traditional, legacy cloud. Start deploying your services on InstaCloud today.
DeepSeek 在 GitHub 开源 DeepGEMM-Ascend,一个面向华为昇腾 NPU 的简洁高效矩阵乘法算子库。该仓库定位为 Ascend NPU 上的 GEMM 内核实现,延续 DeepGEMM 的轻量高效路线。
HPE 提出企业 AI 从按 token 消费转向自建容量的判断框架:当需求稳定、可预测且规模足够时,拥有算力可能比逐次购买更经济。文中引用 Deloitte 2026 企业 AI 状况报告称,2025 年员工 AI 使用率上升 5%,至少 40% 的 AI 项目进入生产的公司比例预计半年内翻倍。企业需先回答三个问题:需求是否稳定、在什么使用水平下自建更划算、能否通过采用与治理让容量保持高产。
vLLM 官方博客发布分离式推理(disaggregated serving)实战指南,讲解如何将 prefill 与 decode 拆分为独立实例、把 tokenization 和解析移到无 GPU 的 /render 与 /derender 前端。
xAI 的 Grok 4.7 已在 Amazon Bedrock 上线,提供 500K token 上下文窗口和 low、medium、high、xhigh 四档可配置推理强度,通过 bedrock-runtime 端点的跨区域推理配置文件提供服务,支持 Responses、Chat Completions 和 Converse API。
推荐理由:梳理了 Grok 4.7 在 Bedrock 上的接入方式、推理档位与成本取舍,便于评估长任务智能体的落地配置。
据知情人士消息,AI 推理基础设施服务商 Modal Labs 正接近完成由 Accel 领投的 7.5 亿美元融资,估值达 157.5 亿美元,较四个月前 3.55 亿美元融资时的 46.5 亿美元估值增长两倍多。
Databricks 介绍其内部新模型发布流程:通过 Unity Gateway 让全体员工首日获得 Opus 5.5、GPT-6 Sol 和 GPT-Luna 的实验性访问,并用四类按用户预算控制成本。
推荐理由:Databricks 以自身一万多名员工的实际 rollout 数据为例,给出一套可复用的新模型评估与推广方法,含具体成本对比。
Databricks 发布 Lakebase Search,通过 lakebase_vector 和 lakebase_text 两个扩展为 Lakebase Postgres 带来可扩展向量搜索与 BM25 全文检索,已在 AWS 和 Azure 正式可用。
Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together, we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. https://nvda.ws/4hcoq7m
SemiAnalysis 分析指出,稀疏注意力只在 SDPA 运算中降低 KV cache 的显存与带宽需求,并不减少整体显存容量占用,因为 top-k 选择仍需完整上下文驻留 HBM。
AWS 宣布 Claude Sonnet 5.5 在 Amazon Bedrock 和 Claude Platform on AWS 上线,定位为更高效、单任务成本更低的 Sonnet 模型,适合范围明确的编码与知识工作。
推荐理由:文章给出 Sonnet 5.5 在 Bedrock 上的能力定位与调用方式,读者可据此判断它适合承接哪类持续运行的编码任务。
Vinext 1.0 从 AI 实验毕业为生产就绪框架,让开发者可以在 Vite 上运行 Next.js 应用。该版本带来高级缓存预热、更广泛的兼容性以及自动化测试流水线。
Vinext 1.0 graduates from an AI experiment to a production-ready framework, letting developers run Next.js apps on Vite. This release brings advanced cache warming, broader compatibility, and an automated testing pipeline. https://cfl.re/4dcK6hk #BirthdayWeek
AWS 发布教程,演示如何用 vLLM-Omni DLC 在 Amazon SageMaker AI 上部署 Qwen3-TTS,通过 SageMaker 双向流式连接实现文本流入、24 kHz PCM 音频分块流出,并用 Gradio 应用验证。
AWS 发布 vLLM-Omni 图像视频生成示例,在同一 vLLM-Omni DLC 上部署两个 SageMaker AI 端点:实时端点跑 FLUX.2-klein-4B 文生图,异步端点跑 Wan2.1-VACE-1.3B 图生视频。
Mistral 在慕尼黑开设德国中心,组建专注 Physics AI 与工业 AI 的研究团队,并计划到 2030 年建成 1 吉瓦欧洲算力。该中心将携手 BMW 开展碰撞仿真与工程 AI 合作、与 Siemens Energy 推进工业 AI 应用,并与慕尼黑工业大学(TUM)合作利用风洞设施开发汽车空气动力学数字孪生。
AWS 博客介绍用 Amazon Nova Act 配合 Amazon Bedrock AgentCore 构建智能体驱动的合成监控方案,以自然语言动作替代 Selenium、Playwright 的 DOM 选择器脚本。
AWS 发布跨账户自动化管理 Amazon Textract 适配器生命周期的方案,提供 CloudFormation 和 Terraform 模板、跨账户适配器提升流程、生产安全配置及文档预分类路由模式。
Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together, we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. https://nvda.ws/4hcoq7m
NVIDIA DSX MaxLPS 通过策略管控的电力共享,在参与资源之间动态分配电力,让客户在正常运行时不再为 GPU 峰值功耗预留保护性缓冲。AI 工厂通常按所有 GPU 同时达到峰值功耗的极端情况配置电力,导致大量基础设施在平时被闲置。该方案旨在提升 AI 工厂的吞吐量与效率。
一样。@useblacksmith 一直是超棒的赞助商,但我们得分散负载。 我的计划是让 codex 决定哪些测试真正需要跑,大幅削减 CI,改为每小时跑一次测试。
CI has become the top bottleneck of every engineering team I talk to (including Lindy). Our CI spend has become stratospheric.