推理中的计算与数据搬运:MoE 如何改变服务结构与成本
SemiAnalysis 长文拆解 MoE 推理的计算与数据搬运,指出 MoE 不只增加参数量,更改变了每个 token 激活哪些张量、哪些数据必须就近放置,以及内存搬运、存储与调度如何影响有效吞吐。
SemiAnalysis 长文拆解 MoE 推理的计算与数据搬运,指出 MoE 不只增加参数量,更改变了每个 token 激活哪些张量、哪些数据必须就近放置,以及内存搬运、存储与调度如何影响有效吞吐。
微软发布 Microsoft Agent 365 分步安全配置指南,覆盖智能体从创建到身份、数据保护与运行时监控的完整流程。指南涉及 Entra、Purview 和 Defender 三大平台,用于配置智能体身份管理与运行期防护。
🔥社区从不停下折腾的脚步。 不只是基于 H3 做开发,还在不断深入内部,寻找让它更聪明的新方法。
流行のJevをMiniMax H3に組み込んで、動画生成を高速化してみた!Attention処理のスパース化にJevを使用。 ・層ごとにJevが重要度を判定(4step 49層が対象) ・Jevがスパース率1%, 3%, 5%, 10%を選択 RTX4070で6分7秒→3分34秒で41.7%短縮!動画生成中にJevクラウドに問合せしているのに速い!
vLLM 公布 Qwen3.8-2.4T 在 GB300 NVL72 集群上的 PD 分离部署结果:8K/1K 负载下高吞吐场景达每 GPU 5000 total token 吞吐,低延迟场景每用户 180 生成 token,并给出完整 pareto 前沿。
Hugging Face 发布 tokenizers v1,输出与 v0.23 完全一致的 token ID,但速度常达 v0.23 的数十倍。v1 将单个 crate 拆为 workspace,引入无分配合并、bitcannon SIMD 分片、侵入式双向链表合并循环、线程本地词缓存和原生多线程并行。
NVIDIA 发布 AIPerf,用于对大规模 LLM 推理做基准测试,解决 curl 命令、手写 asyncio 脚本或一次性压测工具受单进程性能、Python GIL 并发上限等问题。该工具面向模型部署后的性能评估场景。
この3x3画像を1枚 r2v でAI動画化すると以下の設定でおおよそそのまま映像になる Minimax H3 Max r2v 480p 15秒 Prompt調整:Quality 参照強度:標準 「左上から右下のパネルにカットが切り替わる2Dアニメーション、 複数パネル禁止、BGM禁止」
播客好友 @stevenbjohnson 展示了他最喜欢的 Notebook 移动应用用法之一 🤯
The camera feature in the @Gemini_Notebook mobile app is so transformative for on-the-go research. I was up in the Sierras working on a new project, and I just took photos of everything, like this museum display. Then I asked for a detailed report of all the info in the image. The text below is what I got back. (I fact-checked it myself and it was 99% accurate -- and some of the text it transcribed is so blurry in the image that it was hard for me to read it.) Next step is to generate documents like this for all the photos I took, and then ask Notebook to highlight all the information that adds something new to the existing knowledge base of sources I've already collected, or is particularly relevant to my latest writing and outline for the project. Truly magical. Mariposa Museum Exhibit Reference: Mariposa in 1859 & Gold Rush Era This reference document consolidates all text, photographic captions, historical statistics, newspaper clippings, and exhibit overlays displayed on the Mariposa Museum exhibit panel regarding Mariposa during and after the Gold Rush. 1. Exhibit Overview & Key Headlines Main Title: "THIS WAS MARIPOSA IN 1859 – ONLY 10 YEARS AFTER THE GOLD RUSH BEGAN." Historical Context Sub-headline: "PORTIONS OF THE TOWN HAD BEEN REBUILT AFTER THE FIRE OF 1858.. & PARTS OF IT WERE DOOMED TO BE DESTROYED IN 1866" Display Overview: The exhibit centers on a large 1850s panoramic photograph of Mariposa, annotated with street names, landmark locations, and population statistics, flanked by contemporary hotel advertisements, fire reports, medical artifacts, and photographs of civic buildings. 2. Demographic & Real Estate Statistics (1850s vs. Present) MetricHistorical Value (1850s Gold Rush Peak)Modern Value Town Population~3,000 residents~1,800 estimated Entire Mariposa Diggings Area~15,000 residents— Commercial Establishments15 to 20 stores & saloons, plus hotels— Town Lot Prices00 to 00 per lot— 3. Background Photograph & Civic Infrastructure Background Photo Date: Taken in the 1850s, capturing the rapid growth of the settlement following the initial gold strike. 1854 Mariposa County Courthouse: Shown in the background panorama prior to the construction of its iconic clock tower. A separate framed photograph depicts the completed white wood-frame courthouse with a white picket fence and clock tower. Clock Tower History: The clock mechanism was imported from England and is an 8-day, manually wound instrument. It remains operational today, maintained by the Mariposa Public Works Department. Annotated Overlay Locations on Panorama: 1854 Courthouse: Located at the upper edge of town on Jones Street. Jones Street & Bullion Street: Upper residential and civic thoroughfares. Charles Street (Main Street): Primary commercial artery running through the center of the valley floor. Schlageter Hotel Site: Positioned along Main Street. Mariposa Creek: Flowing along the foreground basin of the town diggings. 4. The Fire of 1866: Mariposa's Second Great Conflagration Below is the complete transcript of the Mariposa Gazette report featured on the panel regarding the disaster of August 25, 1866 (following the earlier destructive fire of 1858): Article Text: "FIRE! MARIPOSA'S SECOND GREAT FIRE" A few minutes after 6 p.m. on Saturday, August 25, 1866 Mariposa was again ruined by a disastrous fire (first in 1858). According to the Gazette (the building was damaged but not destroyed), the fire was believed to have started when "a recently imported printer stepped inside the Free Press office and lighted a cigar. The match had evidently been dropped carelessly amongst the papers on the floor. The Free Press office was located near the corner of Main and 7th Streets and by ten minutes the fire had spread through two blocks. The fire crossed Main Street to the Odd Fellows building and the Methodist Church and soon the buildings on the block between 6th and 7th Streets were burning like so much chaff." Seven full blocks, except for four fire-proof buildings, were totally destroyed. "In one hour about 60 buildings and 77,000 worth of property were destroyed." By early Monday morning the men were clearing away debris and by press time the following Saturday, the Gazette reported that several temporary business structures had already been erected and were open for business. Inventory of Buildings Destroyed in the 1866 Fire: Residential & Civic: 14 Dwellings, 1 Church, 1 Odd Fellows and Masons Hall. Media & Printing: 1 Newspaper Office (Free Press), 1 Newspaper Depot. Hospitality & Retail: 3 Hotels, 5 Retail Stores, 1 Saddlery Shop, 9 Liquor Saloons (several equipped with billiard tables). Services & Trades: 2 Livery Stables, 3 Law Offices, 1 Drug Store, 3 Blacksmith Shops, 2 Carpenter Shops, 2 Shoemaker Shops, 1 Tailor Shop, 2 Butchering Establishments. Outbuildings: Numerous outhouses, private stables, and auxiliary structures. 5. Commercial Hotels & Lodging Gallison Hotel (1887 Advertisement) Location: Main Street, Mariposa (center of business district, opposite Odd Fellows' Hall). Proprietor: Winslow Gallison. Management: Mrs. Gallison individually superintended all internal departments of the hotel. Amenities: Newly furnished rooms, first-class table dining. Mariposa Hotel (1887 Card) Location: Corner of Main and Fifth Streets. Proprietor: Charles A. Schlageter. Target Market: Accommodated general travelers as well as Yosemite tourists on short notice. Amenities: Family rooms, well-lighted parlors, good table, and bath facilities. The Schlageter Hotel Date Built: Built in the 1850s. Architecture: Prominent two-story wooden structure featuring full upper and lower covered verandas. 6. Medical Artifacts & 19th-Century Therapeutics Old Mariposa Hospital A framed historical photograph depicts a two-story wood-frame hospital building with a prominent front porch and side wing (annotated "from Chic Allingham"). Dr. D. Jayne's Family Medicines (1880 Display Broadside) 1. Jayne's Specific for Tape-Worm Diagnosis: Describes tapeworm infections as widespread, noting that discharging white or yellowish segments ("resembling gourd seeds") is the only positive diagnostic proof. Pricing & Ordering: .00 per dose, shipped nationwide via mail from 242 Chestnut Street, Philadelphia. Usage Instructions: Dissolve powder in a pint of boiling water, drink in three equal hourly doses on an empty stomach. Follow with Cathartic medicines if bowels do not operate in three hours. 2. Dr. D. Jayne's Sanative Pills Formulation: Concentrated, sugar-coated pills. Sold in 50-pill boxes (-bash.25) or 15-pill specimen packets (-bash.10). Prescribed Ailments: Advertised for liver complaints, gout, jaundice, dyspepsia, rheumatism, kidney affections, fevers, nervousness, skin diseases, melancholy, sick headache, and costiveness (constipation). 3. Exhibit Commentary Note A small museum card mounted below the broadside reads: "A man advertises for 'a competent person to undertake the sale of a new medicine' and adds innocently 'It will prove profitable to the undertaker.'"
Lucius AI 用 AlloyDB for PostgreSQL 承载覆盖五大洲的招标平台,将语义搜索迁移到 ScaNN 索引后,代表性生产查询延迟从 1.14 秒降至 24 毫秒,提速 47 倍。
GitHub 用 GitHub Copilot app 和 Copilot CLI 把 Copilot agent runtime 从 TypeScript/Node.js 完全重写为超过 80 万行生产级 Rust,AI 智能体编写了大部分代码,跨 128 个 PR 增量合入 main,性能提升数个数量级,主要由一名开发者几个月内完成。
推荐理由:GitHub Copilot 运行时迁移 Rust 的完整复盘,给出智能体并行协作、提示缓存与评审流程的可迁移工程方法。
NVIDIA 展示了一套智能体 AI 工作流,用于为物理 AI 系统准备和验证数字孪生。智能体可检查 3D 场景、在 OpenUSD 中编写仿真相关数据、添加物理属性、渲染预检视图,并对照 SimReady 要求验证结果。该流程覆盖从 Blender 场景到面向 NVIDIA 的仿真就绪 OpenUSD 交付。
Midjourney 每周办公时间 - 9/16 https://x.com/i/spaces/1qJVmyRpDbYGB
NVIDIA 技术博客介绍 cuTile Rust(cutile-rs),一个用 Rust 编写 GPU kernel 的 tile 系统,将 Rust 所有权模型扩展到 tile 级 GPU kernel,把可变输出拆分为互不重叠的片段,并在 kernel 启动间保持主机侧所有权契约。
OpenAI 介绍如何通过 ChatGPT Work 和 Codex 的分析功能,帮助团队了解 AI 使用量与支出、识别培训需求,并将采用情况与业务成果关联。
Together AI 发布从闭源模型迁移到开源模型的指南,称采用托管服务可将迁移周期从数月到数年缩短为数周到数月。方法分发现、评估、适配、决策、生产五步,核心是用真实流量回放而非通用基准做评估,并按系统提示词、推理参数、上下文工程、微调四个杠杆迭代适配;文中提到部分客户迁移后成本最多降低 70%,可用 10% 流量的金丝雀部署开始上线。
NVIDIA 技术博客对比 Dense 与 MoE 两种模型架构,说明参数组织方式对性能的影响。以 Nemotron 3.5 Lightning 为例,该模型总参数 30B,但每个 token 仅激活 3B 参数,依靠 MoE 架构按 token 选择部分参数,从而在保留大模型容量的同时降低单次计算量。文章围绕活跃参数、吞吐量与选型时机展开分析。
IBM Research 与 Hugging Face 在 ALTK-Evolve 中推出一致性指南和 Consistency Analyzer,用于诊断和改善智能体重复运行的不稳定性。
推荐理由:原文给出一致性差距的量化诊断方法与开源实现,读者可据此评估和改进智能体在重复运行下的可靠性。
NVIDIA 技术博客介绍如何用 NVIDIA FLARE 将联邦学习从单服务器、少量客户端的简单部署扩展到跨 Docker、Kubernetes 和 Slurm 的共享基础设施。随着项目规模增长,挑战从运行算法转向运营共享基础设施:按需分配 GPU、隔离多个研究任务,并让每个参与机构保留对自身数据的控制权。
NVIDIA 技术博客介绍如何用 NVIDIA Transformer Engine 在 JAX 中加速 Dropless MoE 训练。MoE 通过条件计算实现高效训练,DeepSeek、Qwen、Mixtral 等模型以远低于稠密模型的训练算力达到或超越其性能。文章针对传统 MoE 依赖共享稠密 FFN 的做法,给出 Dropless 训练路径。
vime 联合 RL-Kernel 在 AMD Instinct MI300X 上实现训练与 rollout 的逐位数值一致性,8× MI300X 跑 Qwen3-8B GRPO 实验连续 200 步 mismatch_count = 0、max_abs_diff = 0。
vLLM 通过自适应投机 token 预算、KDA 前缀检查点、零拷贝混合 KDA 批次和延迟 MXFP4 收尾等优化,将 Kimi K3 服务性能从 v0.27.1 提升至 main:延迟降低 56%–60%,吞吐量提升 2.2–2.8 倍,TTFT 降低 72%–85%(并发 1、4、16,8K/1K 负载,TP8)。
GitHub 日本和韩国营销负责人把活动运营流程交给 GitHub Copilot,从单个 GitHub Issue 出发自动完成落地页复制、UTM 链接生成、报名名单清洗和会后报告。
掌握 AI 工程技能,你就能主动塑造构建过程:影响要构建什么,并驱动构建循环。以下是实现这一点的关键技能。https://x.com/i/article/2098450134883594240
Nathan Lambert 在 Interconnects 发布开源 AI 与开放模型阅读清单,收录近几年他认为是该领域最佳的文章,并称可作为了解该领域现状的全面概览。
GitHub Copilot 应用内置 diff、终端和浏览器三个面板,让开发者无需在编辑器、终端和浏览器之间切换即可完成 AI 编码闭环。diff 面板以绿色标注新增、红色标注删除,支持接受改动、留言或让 Copilot 继续修改;终端面板可直接运行项目命令并支持多窗口;浏览器面板可用 Pick & Polish 工具选中元素并让智能体调整。
NVIDIA 通过全栈 NIM 优化,在 Nemotron 3 Ultra 上实现 2.5 倍并发用户量。该优化针对生产环境部署大语言模型时,在现有 GPU 基础设施上提升并发服务能力并保持交互响应速度的需求,对提示词长、上下文跨步骤复用的智能体 AI 工作负载尤为关键。
Together AI 的 kernels 团队为 NVIDIA Vera Rubin NVL72 平台在 ThunderKittens 中新增了 NVFP4 与 FP8 GEMM 支持。
Hugging Face 博客介绍 TRL v1.14 的 AsyncGRPOTrainer 新支持:只训练 LoRA 适配器并仅同步适配器(rank-1 仅几 MB)到 vLLM,训练器与 vLLM 副本作为独立 HF Jobs 运行在分开的机器上,通过挂载 Storage Bucket 共享适配器路径,无需 NCCL 或共享本地盘。
MiniMax M3 在 AMD Instinct MI355X 上的 vLLM 服务性能大幅提升:SemiAnalysis InferenceX 基准显示,并发 32 时 MXFP8 标准服务从 109.1 升至 342.4 output tokens/s/GPU(3.14×),中位 TTFT 从 1.46 秒降至 0.67 秒。
NVIDIA 技术博客解析了 Encode-Prefill-Decode(EPD)分离架构这一多模态模型推理优化技术,它将视觉编码器阶段与 prefill、decode 阶段分离。该方案对图像密集提示词、中短输出和量化 MoE 模型最有效,配合 NVIDIA Dynamo 可实现最高 5 倍加速。
Mistral 帮助一家欧洲能源运营商将 4 万行 Fortran 77 的油藏模拟器迁移到 C++,该代码库没有测试套件和集中文档。团队先搭建数值对齐校验框架,用自定义解析器生成调用树并借助 Vibe CLI 启动上百个智能体补文档,最终采用人工把关的 coder、tester、reviewer 智能体工作流逐模块迁移。
vLLM 为 GLM 5.3 引入 Hybrid HiSparse 卸载机制,在单台 8× H200 节点上首次实现 100 万上下文长度运行,并在各上下文长度下大幅提升并发量。该机制基于稀疏 MLA 的 top-K 选择,将未被选中的 KV cache 卸载至 CPU,仅在 KV cache 承压时才付出 CPU-GPU 传输代价,热缓冲页与常驻页共用同一 HMA 块池。
Sierra 发布呼叫中心 AI 运营与推广指南,主张把 AI 部署当作运营模式变革而非单点演示,需在迁移流量前设定基线与扩展标准。指南覆盖语音、消息、邮件等渠道,Sierra 支持单一智能体跨渠道部署,Voice 支持呼入呼出、语音模拟、多语言与带上下文升级。落地需围绕路由与队列行为、人力排班、人机质量校准等环节设计运营模型。
NVIDIA 团队用 NemoClaw 构建了一个记忆驱动的"幕僚长"智能体,通过名为 self model 的人类可读知识层保存智能体记忆。该智能体面向企业工作中随时间变化的邮件、决策、项目和待办事项,避免每次启动都需重建上下文。
NVIDIA 发布技术指南,介绍如何在 Jetson 边缘设备上部署和优化具备多步推理能力的模型。此前这类模型体积过大,无法在边缘硬件本地运行,开发者只能将推理请求路由至数据中心,带来网络依赖、成本上升与数据外泄风险。该指南称这一限制正在被打破。
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: http://z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: http://huggingface.co/zai-org/GLM-5.3-Flash API: http://docs.z.ai/guides/llm/glm-5.3-flash Coding Plan: http://z.ai/subscribe ZCode: http://zcode.z.ai/en Chat: http://chat.z.ai AutoClaw: http://autoclaw.z.ai
推荐理由:原文给出本地推理提速的具体配置和硬件门槛,读者可以直接复用到自己的 GLM-5.3-Flash 部署中。
NVIDIA 技术博客介绍如何在联邦 Kubernetes 与 AI 平台间传递用户身份。现代 AI 平台中用户从中央门户进入、打开受治理数据集、启动 notebook 并调用另一集群中的助手,身份在每一步都跨越控制面与数据面边界,传统单点登录在此失效。