跳到正文

#部署/工程

今日 49 条
9月22日周二
9月21日周一
  1. Qwen42

    感谢 @sgl_project 的 day-0 支持!🙌 SGLang-Diffusion 现已支持 Qwen-Image-2.1:文生图、多图编辑,以及透明 RGBA 输出。快来试试!🎨

    引用SGLang@sgl_project

    Day-0 support for @Alibaba_Qwen’s Qwen-Image 2.1 is here in SGLang-Diffusion! 🖥️ Native precision on a single RTX 4090 24GB with CPU offload - 1024×1024 generation in 18.7s and image editing in 21.7s with 22.7 GiB peak GPU memory during requests. - On an RTX PRO 6000 96GB: 8.0s generation and 9.6s editing. 🎨 Text-to-image, multi-image editing, and transparent RGBA output—all with one checkpoint. ⚡ Native inference with TP/SP, LoRA, and OpenAI-compatible APIs. 40 denoising steps, one image per request, warmed HTTP latency including PNG output. No quantization. Cookbook and GPU-specific commands below 👇

9月19日周六
9月18日周五
  1. MiniMax (official)40

    Nunchux AI 与多校研究者推出 VC-Attention,为 MiniMax-H3 带来免训练低比特注意力加速,在 B200 上比 FlashAttention-4 快 1.6×、B300 上快 1.5×,保真度优于 SageAttention2。

    引用Nunchux AI@NunchuxAI

    Introducing VC-Attention: fast and accurate low-bit attention without retraining. On MiniMax-H3, VC-Attention speeds up attention by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. It also works with existing sparse attention methods. Two key innovations: • V-Smooth reduces value quantization error. • ExpCast-FP8 speeds up softmax. Nunchux Attention, our proprietary extension, pushes the speedup to 1.9× on B200 and 1.8× on B300. Blog: http://www.nunchux.ai/blog/attention-is-the-video-bottleneck Technical Report: http://arxiv.org/pdf/2609.15810 Joint work by researchers at MIT, CMU, UC Berkeley, Stanford, and NVIDIA.

9月17日周四
  1. Unsloth AI67

    Unsloth 发布更新的 Unsloth Docker 镜像,可免设置地在本地训练和运行 500+ 模型,提供新 GUI 和 notebooks 两种工作流,依赖已预装,支持 NVIDIA 和 AMD。指南见 https://unsloth.ai/docs/get-started/install/docker ,代码见 https://github.com/unslothai/unsloth 。

    引用Unsloth AI@UnslothAI

    Introducing Unsloth Desktop 🦥 The first desktop app to run and train models locally. • Open-source. Runs on Mac, Windows and Linux • Supports MLX, diffusion image/video, audio, GGUF • Connect Claude Code and Codex to local LLMs • 50% more accurate, self-healing tool calls + sandboxed code exec • Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac • Train models 2× faster with 70% less VRAM • Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF) • Use Unsloth’s OpenAI-compatible API and cloud models • Securely deploy LLMs remotely and access anywhere Unsloth Desktop is now available on http://unsloth.ai and GitHub. GitHub: https://github.com/unslothai/unsloth Blog and Guide: https://unsloth.ai/docs/desktop

    推荐理由:原文给出免配置的本地训练运行入口和适用硬件范围,读者可以据此评估是否替换现有部署流程。

  2. jietang66

    唐杰发文复盘,GLM-5.3-Flash 从首次在国内加速器上运行到承接全部生产流量只用两周,端到端吞吐达 3.2 倍,大量工作由 GLM-5.3 驱动的 Infra Agent 完成。

    引用Z.ai@Zai_org

    We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone. https://z.ai/blog/glm-built-its-inference-infrastructure

    推荐理由:作者复盘了 GLM-5.3 智能体优化推理基础设施的两周过程,提出了可迁移的分层密集反馈方法与工程师角色转变的判断。

  3. GitHub Blog85

    GitHub Copilot 用 Copilot 把运行时迁移到 Rust:80 万行代码、128 个 PR

    GitHub 用 GitHub Copilot app 和 Copilot CLI 把 Copilot agent runtime 从 TypeScript/Node.js 完全重写为超过 80 万行生产级 Rust,AI 智能体编写了大部分代码,跨 128 个 PR 增量合入 main,性能提升数个数量级,主要由一名开发者几个月内完成。

    推荐理由:GitHub Copilot 运行时迁移 Rust 的完整复盘,给出智能体并行协作、提示缓存与评审流程的可迁移工程方法。

9月16日周三
  1. NVIDIA Blog(RSS)53

    Emerald AI、Google 与 NVIDIA 发起 AI Energy Management Alliance 推动灵活用电数据中心

    Emerald AI、Google 与 NVIDIA 宣布发起 AI Energy Management Alliance(AEMA),推动数据中心根据电网状况动态调整用电。联盟主张技术中立、按性能衡量灵活性,制定并网前的响应义务、统一技术要求和更快并网通道,汇聚 AI 平台、数据中心、电力公司与电网运营商等价值链成员,以提升现有电网容量利用、缩短 AI 设施并网时间。

  2. Together AI 研究与产品博客(RSS)53

    Together AI 详解从闭源模型迁移到开源模型的策略

    Together AI 发布从闭源模型迁移到开源模型的指南,称采用托管服务可将迁移周期从数月到数年缩短为数周到数月。方法分发现、评估、适配、决策、生产五步,核心是用真实流量回放而非通用基准做评估,并按系统提示词、推理参数、上下文工程、微调四个杠杆迭代适配;文中提到部分客户迁移后成本最多降低 70%,可用 10% 流量的金丝雀部署开始上线。

  3. Apple Machine Learning Research(RSS)38

    Glyph:面向企业数据目录列描述与敏感本体标注的多策略智能体系统

    Apple 研究团队提出 Glyph,一个将列描述生成与列类型标注建模为有状态图编排的多智能体 LLM 生产系统。其 Descriptor 通过推理-行动工具循环从企业 GitHub 按需检索管道源码来支撑生成,Tagger 并行运行描述、业务线正则与元数据三种策略,并用 RRF 融合排序结果,从 275 叶节点的数据分类本体中打标。

  4. SemiAnalysis 长文 RSS(RSS)67

    SemiAnalysis 反驳数据中心暂停令正在扼杀美国建设潮的说法

    SemiAnalysis 分析认为数据中心暂停令严重拖慢美国建设的说法不准确。其模型预测 2027 年美国新增 38GW IT 容量,是 2026 年的两倍以上;约 300 个地方暂停令中实际被直接延迟的容量仅约 2.3GW,其中纽约州约 0.8GW、地方限制约 1,525MW,主要由俄亥俄 AWS 园区等三个项目构成。

  5. NVIDIA Technical Blog(开发者技术博客 · RSS)21

    NVIDIA Groq 3 LPX 的确定性执行如何在 NVIDIA Vera Rubin 上驱动高能效高交互推理

    NVIDIA Groq 3 LPX 通过确定性执行,在 NVIDIA Vera Rubin 平台上实现高交互推理的能效提升。该方案针对 AI 工厂的功耗约束,以每瓦性能而非原始吞吐量作为衡量 AI 平台价值的核心指标。Vera Rubin 平台正是为在有限功耗预算内最大化输出而设计。

9月15日周二
  1. NVIDIA Technical Blog(开发者技术博客 · RSS)28

    NVIDIA FLARE 如何跨 Docker、Kubernetes 和 Slurm 扩展联邦学习

    NVIDIA 技术博客介绍如何用 NVIDIA FLARE 将联邦学习从单服务器、少量客户端的简单部署扩展到跨 Docker、Kubernetes 和 Slurm 的共享基础设施。随着项目规模增长,挑战从运行算法转向运营共享基础设施:按需分配 GPU、隔离多个研究任务,并让每个参与机构保留对自身数据的控制权。

9月14日周一