Amazon SageMaker HyperPod 与 Qumulo 实现跨区域训练
Amazon SageMaker HyperPod 搭配 Qumulo 的 Cloud Native Qumulo(CNQ)与 Cloud Data Fabric(CDF),可让训练集群直接读取另一 AWS Region 或本地数据中心的数据集,无需复制数据或修改代码。
Amazon SageMaker HyperPod 搭配 Qumulo 的 Cloud Native Qumulo(CNQ)与 Cloud Data Fabric(CDF),可让训练集群直接读取另一 AWS Region 或本地数据中心的数据集,无需复制数据或修改代码。
很多人想知道如何用我们新的编辑模型来实现一致的角色等等——所以我们为你们做了一个视频 <3
AWS 发布 WhisperX Deep Learning Container(DLC),在 OpenAI Whisper 基础上加入 wav2vec2 强制对齐的逐词时间戳和说话人分离,可部署到 Amazon SageMaker AI 实时或异步端点,无需自建镜像。
NVIDIA 技术博客解析了生物基础模型的高效 MoE 训练方案:相比每个 token 都要过全部层的稠密 Transformer,MoE 用多个专家子网络、每个 token 只激活其中一小部分,从而降低训练与推理算力开销。文章指出,随着语言模型规模增长,稠密架构的扩展成本越来越高。
一位工程师在 Databricks 上构建了基于智能体的安全审查层,用 7 个各司其职的智能体承接可重复的审查工作,把人工判断留给高风险和模糊决策。
We made claude.ai 3x faster in two weeks. Here’s how we use Claude to measure, debug and improve performance. Prompts and methods included. https://claude.dev/blog/how-we-made-claude-ai-faster/
推荐理由:Anthropic 团队分享了用 Claude 测量、调试并让 claude.ai 在两周内提速 3x 的具体方法,文中附带了可直接复用的提示词。
NVIDIA 指出 GPU 集群即使所有 GPU、网络链路和 pod 均报告健康,512-GPU 训练任务仍可能性能不达标或失败,原因可能是单个慢速 GPU、负载下性能劣化的链路,或悄悄将流量导向慢速路径的配置。运维人员往往要到训练运行数小时后才会发现问题。
NVIDIA Warp 与 MuJoCo Warp(MJWarp)可将兼容的 MuJoCo 模型扩展到最多 2,048 个并行 GPU 环境,本文以 SO-101 机械臂为例演示从 CPU MuJoCo 工作流迁移到 GPU 批量仿真的过程。
GitHub Copilot 应用重建了 pull request 视图,用一个含 2200 个文件、超 100 万行改动和 400 多条行内评论的开源 PR 做压力测试。其做法是把文档高度拆成确定性的代码几何与动态评论块两套几何:代码行高提前精确计算,评论高度则按需测量、修正幅度小且锚定在用户当前查看位置,从而避免滚动跳动。
智能体在工作前后喜欢大量读取以收集上下文,所以针对读取做优化会带来很大差别! 如果你在构建智能体,我强烈推荐读一读这个(或者把它交给你的智能体,让它去实现这些发现)
We've reduced token costs in Cursor by 7% with no drop in agent quality. Savings came from tighter prompts, selective tool loading, better caching, and compressed file reads.
一篇用可交互示例讲解 shadow roots 的教程,演示了样式封装、继承、slots、parts 以及 JavaScript 访问方式。内容展示 shadow roots 如何创建带有私有样式表和元素的隔离 DOM 树,并以特定且受控的方式与页面 DOM 交互。
Together AI 发布教程,演示如何以 Qwen3.5 4B 为基础模型,用 MultiNLI、BoolQ、Banking77 等 6 个数据源共 37,840 条样本微调一个 Jev 式分类模型,训练成本约 $17.0,耗时约 25 分钟。
肝不动了....GPT的几个和后端/Agent性能测试明天再说吧.....
@karminski3 效率也太高了!你不睡觉的吗😳
如何在 60 秒内在 @CloudflareDev AI Gateway 上配置来自 @typesafeai 的 Jev。
NVIDIA 提出 AI 智能体评估需从单次工具调用打分转向整项任务完成度打分,核心是看智能体能否在真实环境中连续执行数十次工具调用并在某步失败后恢复。仅评估模型回答是否"听起来对",几乎无法反映工作是否真正完成。
SemiAnalysis 长文拆解 MoE 推理的计算与数据搬运,指出 MoE 不只增加参数量,更改变了每个 token 激活哪些张量、哪些数据必须就近放置,以及内存搬运、存储与调度如何影响有效吞吐。
微软发布 Microsoft Agent 365 分步安全配置指南,覆盖智能体从创建到身份、数据保护与运行时监控的完整流程。指南涉及 Entra、Purview 和 Defender 三大平台,用于配置智能体身份管理与运行期防护。
🔥社区从不停下折腾的脚步。 不只是基于 H3 做开发,还在不断深入内部,寻找让它更聪明的新方法。
流行のJevをMiniMax H3に組み込んで、動画生成を高速化してみた!Attention処理のスパース化にJevを使用。 ・層ごとにJevが重要度を判定(4step 49層が対象) ・Jevがスパース率1%, 3%, 5%, 10%を選択 RTX4070で6分7秒→3分34秒で41.7%短縮!動画生成中にJevクラウドに問合せしているのに速い!
vLLM 公布 Qwen3.8-2.4T 在 GB300 NVL72 集群上的 PD 分离部署结果:8K/1K 负载下高吞吐场景达每 GPU 5000 total token 吞吐,低延迟场景每用户 180 生成 token,并给出完整 pareto 前沿。
Hugging Face 发布 tokenizers v1,输出与 v0.23 完全一致的 token ID,但速度常达 v0.23 的数十倍。v1 将单个 crate 拆为 workspace,引入无分配合并、bitcannon SIMD 分片、侵入式双向链表合并循环、线程本地词缓存和原生多线程并行。
NVIDIA 发布 AIPerf,用于对大规模 LLM 推理做基准测试,解决 curl 命令、手写 asyncio 脚本或一次性压测工具受单进程性能、Python GIL 并发上限等问题。该工具面向模型部署后的性能评估场景。
播客好友 @stevenbjohnson 展示了他最喜欢的 Notebook 移动应用用法之一 🤯
The camera feature in the @Gemini_Notebook mobile app is so transformative for on-the-go research. I was up in the Sierras working on a new project, and I just took photos of everything, like this museum display. Then I asked for a detailed report of all the info in the image. The text below is what I got back. (I fact-checked it myself and it was 99% accurate -- and some of the text it transcribed is so blurry in the image that it was hard for me to read it.) Next step is to generate documents like this for all the photos I took, and then ask Notebook to highlight all the information that adds something new to the existing knowledge base of sources I've already collected, or is particularly relevant to my latest writing and outline for the project. Truly magical. Mariposa Museum Exhibit Reference: Mariposa in 1859 & Gold Rush Era This reference document consolidates all text, photographic captions, historical statistics, newspaper clippings, and exhibit overlays displayed on the Mariposa Museum exhibit panel regarding Mariposa during and after the Gold Rush. 1. Exhibit Overview & Key Headlines Main Title: "THIS WAS MARIPOSA IN 1859 – ONLY 10 YEARS AFTER THE GOLD RUSH BEGAN." Historical Context Sub-headline: "PORTIONS OF THE TOWN HAD BEEN REBUILT AFTER THE FIRE OF 1858.. & PARTS OF IT WERE DOOMED TO BE DESTROYED IN 1866" Display Overview: The exhibit centers on a large 1850s panoramic photograph of Mariposa, annotated with street names, landmark locations, and population statistics, flanked by contemporary hotel advertisements, fire reports, medical artifacts, and photographs of civic buildings. 2. Demographic & Real Estate Statistics (1850s vs. Present) MetricHistorical Value (1850s Gold Rush Peak)Modern Value Town Population~3,000 residents~1,800 estimated Entire Mariposa Diggings Area~15,000 residents— Commercial Establishments15 to 20 stores & saloons, plus hotels— Town Lot Prices00 to 00 per lot— 3. Background Photograph & Civic Infrastructure Background Photo Date: Taken in the 1850s, capturing the rapid growth of the settlement following the initial gold strike. 1854 Mariposa County Courthouse: Shown in the background panorama prior to the construction of its iconic clock tower. A separate framed photograph depicts the completed white wood-frame courthouse with a white picket fence and clock tower. Clock Tower History: The clock mechanism was imported from England and is an 8-day, manually wound instrument. It remains operational today, maintained by the Mariposa Public Works Department. Annotated Overlay Locations on Panorama: 1854 Courthouse: Located at the upper edge of town on Jones Street. Jones Street & Bullion Street: Upper residential and civic thoroughfares. Charles Street (Main Street): Primary commercial artery running through the center of the valley floor. Schlageter Hotel Site: Positioned along Main Street. Mariposa Creek: Flowing along the foreground basin of the town diggings. 4. The Fire of 1866: Mariposa's Second Great Conflagration Below is the complete transcript of the Mariposa Gazette report featured on the panel regarding the disaster of August 25, 1866 (following the earlier destructive fire of 1858): Article Text: "FIRE! MARIPOSA'S SECOND GREAT FIRE" A few minutes after 6 p.m. on Saturday, August 25, 1866 Mariposa was again ruined by a disastrous fire (first in 1858). According to the Gazette (the building was damaged but not destroyed), the fire was believed to have started when "a recently imported printer stepped inside the Free Press office and lighted a cigar. The match had evidently been dropped carelessly amongst the papers on the floor. The Free Press office was located near the corner of Main and 7th Streets and by ten minutes the fire had spread through two blocks. The fire crossed Main Street to the Odd Fellows building and the Methodist Church and soon the buildings on the block between 6th and 7th Streets were burning like so much chaff." Seven full blocks, except for four fire-proof buildings, were totally destroyed. "In one hour about 60 buildings and 77,000 worth of property were destroyed." By early Monday morning the men were clearing away debris and by press time the following Saturday, the Gazette reported that several temporary business structures had already been erected and were open for business. Inventory of Buildings Destroyed in the 1866 Fire: Residential & Civic: 14 Dwellings, 1 Church, 1 Odd Fellows and Masons Hall. Media & Printing: 1 Newspaper Office (Free Press), 1 Newspaper Depot. Hospitality & Retail: 3 Hotels, 5 Retail Stores, 1 Saddlery Shop, 9 Liquor Saloons (several equipped with billiard tables). Services & Trades: 2 Livery Stables, 3 Law Offices, 1 Drug Store, 3 Blacksmith Shops, 2 Carpenter Shops, 2 Shoemaker Shops, 1 Tailor Shop, 2 Butchering Establishments. Outbuildings: Numerous outhouses, private stables, and auxiliary structures. 5. Commercial Hotels & Lodging Gallison Hotel (1887 Advertisement) Location: Main Street, Mariposa (center of business district, opposite Odd Fellows' Hall). Proprietor: Winslow Gallison. Management: Mrs. Gallison individually superintended all internal departments of the hotel. Amenities: Newly furnished rooms, first-class table dining. Mariposa Hotel (1887 Card) Location: Corner of Main and Fifth Streets. Proprietor: Charles A. Schlageter. Target Market: Accommodated general travelers as well as Yosemite tourists on short notice. Amenities: Family rooms, well-lighted parlors, good table, and bath facilities. The Schlageter Hotel Date Built: Built in the 1850s. Architecture: Prominent two-story wooden structure featuring full upper and lower covered verandas. 6. Medical Artifacts & 19th-Century Therapeutics Old Mariposa Hospital A framed historical photograph depicts a two-story wood-frame hospital building with a prominent front porch and side wing (annotated "from Chic Allingham"). Dr. D. Jayne's Family Medicines (1880 Display Broadside) 1. Jayne's Specific for Tape-Worm Diagnosis: Describes tapeworm infections as widespread, noting that discharging white or yellowish segments ("resembling gourd seeds") is the only positive diagnostic proof. Pricing & Ordering: .00 per dose, shipped nationwide via mail from 242 Chestnut Street, Philadelphia. Usage Instructions: Dissolve powder in a pint of boiling water, drink in three equal hourly doses on an empty stomach. Follow with Cathartic medicines if bowels do not operate in three hours. 2. Dr. D. Jayne's Sanative Pills Formulation: Concentrated, sugar-coated pills. Sold in 50-pill boxes (-bash.25) or 15-pill specimen packets (-bash.10). Prescribed Ailments: Advertised for liver complaints, gout, jaundice, dyspepsia, rheumatism, kidney affections, fevers, nervousness, skin diseases, melancholy, sick headache, and costiveness (constipation). 3. Exhibit Commentary Note A small museum card mounted below the broadside reads: "A man advertises for 'a competent person to undertake the sale of a new medicine' and adds innocently 'It will prove profitable to the undertaker.'"
Lucius AI 用 AlloyDB for PostgreSQL 承载覆盖五大洲的招标平台,将语义搜索迁移到 ScaNN 索引后,代表性生产查询延迟从 1.14 秒降至 24 毫秒,提速 47 倍。
Pragmatic Engineer 发布与 Matt Pocock 的播客访谈,讨论其从声乐教师转型开发者与教育者、创办 Total TypeScript(总销售额超 250 万美元)的经历,以及为 AI 编码 Agent 构建技能的实践。
唐杰发文复盘,GLM-5.3-Flash 从首次在国内加速器上运行到承接全部生产流量只用两周,端到端吞吐达 3.2 倍,大量工作由 GLM-5.3 驱动的 Infra Agent 完成。
We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone. https://z.ai/blog/glm-built-its-inference-infrastructure
推荐理由:作者复盘了 GLM-5.3 智能体优化推理基础设施的两周过程,提出了可迁移的分层密集反馈方法与工程师角色转变的判断。
GitHub 用 GitHub Copilot app 和 Copilot CLI 把 Copilot agent runtime 从 TypeScript/Node.js 完全重写为超过 80 万行生产级 Rust,AI 智能体编写了大部分代码,跨 128 个 PR 增量合入 main,性能提升数个数量级,主要由一名开发者几个月内完成。
推荐理由:GitHub Copilot 运行时迁移 Rust 的完整复盘,给出智能体并行协作、提示缓存与评审流程的可迁移工程方法。
NVIDIA 展示了一套智能体 AI 工作流,用于为物理 AI 系统准备和验证数字孪生。智能体可检查 3D 场景、在 OpenUSD 中编写仿真相关数据、添加物理属性、渲染预检视图,并对照 SimReady 要求验证结果。该流程覆盖从 Blender 场景到面向 NVIDIA 的仿真就绪 OpenUSD 交付。
Midjourney 每周办公时间 - 9/16 https://x.com/i/spaces/1qJVmyRpDbYGB
NVIDIA 技术博客介绍 cuTile Rust(cutile-rs),一个用 Rust 编写 GPU kernel 的 tile 系统,将 Rust 所有权模型扩展到 tile 级 GPU kernel,把可变输出拆分为互不重叠的片段,并在 kernel 启动间保持主机侧所有权契约。
OpenAI 介绍如何通过 ChatGPT Work 和 Codex 的分析功能,帮助团队了解 AI 使用量与支出、识别培训需求,并将采用情况与业务成果关联。
Together AI 发布从闭源模型迁移到开源模型的指南,称采用托管服务可将迁移周期从数月到数年缩短为数周到数月。方法分发现、评估、适配、决策、生产五步,核心是用真实流量回放而非通用基准做评估,并按系统提示词、推理参数、上下文工程、微调四个杠杆迭代适配;文中提到部分客户迁移后成本最多降低 70%,可用 10% 流量的金丝雀部署开始上线。
NVIDIA 技术博客对比 Dense 与 MoE 两种模型架构,说明参数组织方式对性能的影响。以 Nemotron 3.5 Lightning 为例,该模型总参数 30B,但每个 token 仅激活 3B 参数,依靠 MoE 架构按 token 选择部分参数,从而在保留大模型容量的同时降低单次计算量。文章围绕活跃参数、吞吐量与选型时机展开分析。
NVIDIA 技术博客介绍如何用 NVIDIA FLARE 将联邦学习从单服务器、少量客户端的简单部署扩展到跨 Docker、Kubernetes 和 Slurm 的共享基础设施。随着项目规模增长,挑战从运行算法转向运营共享基础设施:按需分配 GPU、隔离多个研究任务,并让每个参与机构保留对自身数据的控制权。
NVIDIA 技术博客介绍如何用 NVIDIA Transformer Engine 在 JAX 中加速 Dropless MoE 训练。MoE 通过条件计算实现高效训练,DeepSeek、Qwen、Mixtral 等模型以远低于稠密模型的训练算力达到或超越其性能。文章针对传统 MoE 依赖共享稠密 FFN 的做法,给出 Dropless 训练路径。
vime 联合 RL-Kernel 在 AMD Instinct MI300X 上实现训练与 rollout 的逐位数值一致性,8× MI300X 跑 Qwen3-8B GRPO 实验连续 200 步 mismatch_count = 0、max_abs_diff = 0。
vLLM 通过自适应投机 token 预算、KDA 前缀检查点、零拷贝混合 KDA 批次和延迟 MXFP4 收尾等优化,将 Kimi K3 服务性能从 v0.27.1 提升至 main:延迟降低 56%–60%,吞吐量提升 2.2–2.8 倍,TTFT 降低 72%–85%(并发 1、4、16,8K/1K 负载,TP8)。