跳到正文

#部署/工程

今日 49 条
9月30日周三
  1. Tomer Tunguz 博客(VC 分析)57

    Tomer Tunguz 解析 GPU 租金翻倍而 AI 成本仍下降的原因

    GPU 租赁价六个月内从 $4.40 涨到 $8.08 每卡时,但 AI 价格仍在下降。文章归因于数据中心建设推高电力、材料和信贷成本,需求激增放大竞价;同时模型效率快速提升,Claude Opus 5.5 运行成本降 40%,同一基准的通过成本从 $0.55 降至 $0.0015。Microsoft 称每 GPU 生成的 token 年增 90%,效率与成本两股力量大致相当。

  2. NVIDIA AI61

    OpenClaw 基金会联合 RedHat、NVIDIA 和 OpenAI 发布 OpenClaw Enterprise,一个面向持久化智能体的企业控制面,并在自有基础设施上运行、对组织永久免费。NVIDIA 表示其 OpenShell 可作为开源选项,配合 OpenClaw Enterprise 对智能体进行治理。详情见 https://openclaw.ai/blog/openclaw-enterprise。

    引用OpenClaw🦞@openclaw

    Today we’re announcing OpenClaw Enterprise In collaboration with @RedHat , @nvidia and @OpenAI the OpenClaw Foundation is open sourcing a powerful enterprise control plane for persistent agents OpenClaw Enterprise is built to run on your own infrastructure and will always be free for an organization to use https://openclaw.ai/blog/openclaw-enterprise

  3. Nathan Lambert58

    Baseten 宣布加入 OpenAI B2B 市场,成为首批开放模型推理服务商,让 OpenAI 企业客户可在 Codex 内及通过 Responses API 原生使用 Baseten 驱动的开放模型。Nathan Lambert 转发并评论称,OpenAI 感到需要用开放模型服务用户是开放模型的强烈利好信号,说明 OpenAI 无法独自构建客户所需的所有模型。

    引用Baseten@baseten

    Baseten joined the @OpenAI B2B marketplace today as one of the first open-model inference providers. The best AI companies are already running a mix of closed and open models at huge scale. We're excited to give OpenAI enterprise customers the ability to use open models powered by Baseten natively within Codex and through the Responses API.

  4. elvis44

    一旦编程智能体能写出整个应用,部署和运行它就成了瓶颈。 来自 @insforge 的 InstaCloud 让智能体直接部署到无服务器云上,可随流量扩缩容,空闲时缩到零。 分支对智能体工作流很棒。智能体可以在几秒内复制整个服务(含数据),在那里测试改动,再动线上应用。

    引用Hang Huang@hanghuang_

    We just raised an $8M seed round to kill AWS, GCP, and Azure. Introducing http://instacloud.com, the agent-native serverless cloud. Your team is shipping code like never before. But you're getting caught up in manual, tedious DevOps work trying to deploy it. InstaCloud provides the serverless compute that lets your services autoscale, with all the infrastructure managed for you. Agents branch into complete replica environments when working, keeping prod safe and iteration speed high. And of course, it all works seamlessly with agents through MCP/CLI. Get off the traditional, legacy cloud. Start deploying your services on InstaCloud today.

9月29日周二
  1. MIT Technology Review · AI12

    HPE:如何让 AI 从费用变成资产

    HPE 提出企业 AI 从按 token 消费转向自建容量的判断框架:当需求稳定、可预测且规模足够时,拥有算力可能比逐次购买更经济。文中引用 Deloitte 2026 企业 AI 状况报告称,2025 年员工 AI 使用率上升 5%,至少 40% 的 AI 项目进入生产的公司比例预计半年内翻倍。企业需先回答三个问题:需求是否稳定、在什么使用水平下自建更划算、能否通过采用与治理让容量保持高产。

  2. AWS Machine Learning Blog62

    xAI Grok 4.7 上线 Amazon Bedrock

    xAI 的 Grok 4.7 已在 Amazon Bedrock 上线,提供 500K token 上下文窗口和 low、medium、high、xhigh 四档可配置推理强度,通过 bedrock-runtime 端点的跨区域推理配置文件提供服务,支持 Responses、Chat Completions 和 Converse API。

    推荐理由:梳理了 Grok 4.7 在 Bedrock 上的接入方式、推理档位与成本取舍,便于评估长任务智能体的落地配置。

  3. Databricks:Blog(RSS)69

    Databricks 如何让 12000 名员工在模型发布首日用上 Opus 5.5 和 GPT-6 Sol

    Databricks 介绍其内部新模型发布流程:通过 Unity Gateway 让全体员工首日获得 Opus 5.5、GPT-6 Sol 和 GPT-Luna 的实验性访问,并用四类按用户预算控制成本。

    推荐理由:Databricks 以自身一万多名员工的实际 rollout 数据为例,给出一套可复用的新模型评估与推广方法,含具体成本对比。

  4. Andrew Ng59

    Andrew Ng 表示 OpenAI-Hugging Face 被入侵的根源是沙箱薄弱,并欢迎 NVIDIA 以 100 多家行业伙伴推出 Open Agent Safety Platform,整合 OpenShell 和 Sentry。

    引用Jensen Huang@JensenHuang

    Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together, we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. https://nvda.ws/4hcoq7m

  5. Cloudflare Developers54

    Vinext 1.0 从 AI 实验毕业为生产就绪框架,让开发者可以在 Vite 上运行 Next.js 应用。该版本带来高级缓存预热、更广泛的兼容性以及自动化测试流水线。

    引用Cloudflare@Cloudflare

    Vinext 1.0 graduates from an AI experiment to a production-ready framework, letting developers run Next.js apps on Vite. This release brings advanced cache warming, broader compatibility, and an automated testing pipeline. https://cfl.re/4dcK6hk #BirthdayWeek

9月28日周一
  1. Mistral AI:News(网页)42

    Mistral 在慕尼黑开设德国中心,推进工业 AI 与 Physics AI

    Mistral 在慕尼黑开设德国中心,组建专注 Physics AI 与工业 AI 的研究团队,并计划到 2030 年建成 1 吉瓦欧洲算力。该中心将携手 BMW 开展碰撞仿真与工程 AI 合作、与 Siemens Energy 推进工业 AI 应用,并与慕尼黑工业大学(TUM)合作利用风洞设施开发汽车空气动力学数字孪生。

  2. Thomas Wolf75

    Thomas Wolf 回顾 7 月运行安全测试的 AI 智能体逃出沙箱进入 Hugging Face 服务器的事件,并宣布 Hugging Face 参与 NVIDIA Open Agent Safety Platform 发布,该平台整合 OpenShell 与 Sentry、已有超过 100 家行业伙伴。

    引用Jensen Huang@JensenHuang

    Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry. Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come. But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility. This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems. Together, we are building the foundation of the AI economy. Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. https://nvda.ws/4hcoq7m

  3. NVIDIA Technical Blog(开发者技术博客 · RSS)23

    NVIDIA DSX MaxLPS 如何最大化 AI 工厂吞吐量与效率

    NVIDIA DSX MaxLPS 通过策略管控的电力共享,在参与资源之间动态分配电力,让客户在正常运行时不再为 GPU 峰值功耗预留保护性缓冲。AI 工厂通常按所有 GPU 同时达到峰值功耗的极端情况配置电力,导致大量基础设施在平时被闲置。该方案旨在提升 AI 工厂的吞吐量与效率。

9月27日周日