Hugging Face 发布 AIFS-single-2.0 全 GPU 运行补丁与教程
Hugging Face 发布 AIFS-single-2.0 的补丁与教程,支持通过 Hugging Face jobs 或本地环境在各类 GPU 上运行该模型。仓库名为 huggingface/AIFS-single-2.0-on-all-GPUs,内容为运行补丁与操作教程。
Hugging Face 发布 AIFS-single-2.0 的补丁与教程,支持通过 Hugging Face jobs 或本地环境在各类 GPU 上运行该模型。仓库名为 huggingface/AIFS-single-2.0-on-all-GPUs,内容为运行补丁与操作教程。
Pollen Robotics 发布开源低成本系统 Grabette,用手持夹持器录制约 490€ 的演示数据,无需机器人或遥操作设备,即可自动转成 LeRobot 格式的机器人可用数据集。
Moonshot AI 于 7 月 16 日发布旗舰模型 Kimi K3,为 2.8T 参数 MoE 模型,权重定于 7 月 27 日开放,在 Vals AI 指数排第 2、Artificial Analysis 智能指数排第 3、Frontend Code Arena 排第 1。
推荐理由:作者基于榜单和架构细节分析 K3 对开源模型经济与中美竞争格局的影响,提供了可参考的行业判断框架。
Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. https://thinkingmachines.ai/news/introducing-inkling/ Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. https://thinkingmachines.ai/news/introducing-inkling/ Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
推荐理由:Thinking Machines 发布首个模型 Inkling,开放全部权重并支持 Tinker 微调,可关注其跨模态推理的落地方式。
推荐理由:原文给出多模态能力、开放权重和两个可用入口,读者可以据此了解获取和试用方式。
Thinking Machines Lab 发布多模态 MoE 模型 Inkling,Together AI 在发布当天通过 Serverless 提供托管推理,支持 1M 上下文窗口和 OpenAI 兼容 API。
推荐理由:原文给出 Inkling 的架构细节、参数规模和初步评测数据,读者可以据此判断其推理与多模态能力是否适合接入。
vLLM 宣布 Day-0 支持 Thinking Machines Lab 的 1T 参数多模态模型 TML Inkling,提供 thinkingmachines/Inkling-NVFP4 和 BF16 两个版本。
推荐理由:官方详解了 Day-0 支持背后的 sconv 缓存、TP 分片和 MTP 处理等优化,读者可迁移到类似新架构的推理部署。
OpenAI 在 GitHub 发布 openai/codex-security 仓库,提供用于查找、验证和修复安全漏洞的 Codex Security CLI 和 TypeScript SDK。npm 包名为 @openai/codex-security。
Nathan Lambert 撰文称,白宫正讨论通过新行政令管理开放权重模型,最可能的行动是禁止或无限期推迟能力超过 GPT 5.5、Claude Opus 4.8 或 GLM-5.2 水平的开源模型,且这一时点可能在 6 个月内到来。
推荐理由:作者基于白宫行政令讨论和模型许可会议细节,分析了开源模型禁令的时间线、知识蒸馏争议及其政策走向。
蚂蚁 inclusionAI 在 GitHub 新建仓库 AKernel,定位为面向 Agents 的可编程数据中心级基础设施。该仓库目前仅给出这一句项目描述,未披露模型、参数或性能细节。
蚂蚁 inclusionAI 在 GitHub 上线新仓库 sandboxd,定位为沙箱运行时守护进程(sandbox runtime daemon)。目前仓库仅提供这一句说明,未披露具体功能、支持平台或开源许可等细节。
Thinking Machines Lab exists to empower humanity through advancing collaborative general intelligence. We're building multimodal AI that works with how you naturally interact with the world - through conversation, through sight, through the messy way we collaborate. We're excited that in the next couple months we’ll be able to share our first product, which will include a significant open source component and be useful for researchers and startups developing custom models. Soon, we’ll also share our best science to help the research community better understand frontier AI systems. To accelerate our progress, we’re happy to confirm that we’ve raised $2B led by a16z with participation from NVIDIA, Accel, ServiceNow, CISCO, AMD, Jane Street and more who share our mission. We’re always looking for extraordinary talent that learns by doing, turning research into useful things. We believe AI should serve as an extension of individual agency and, in the spirit of freedom, be distributed as widely and equitably as possible. We hope this vision resonates with those who share our commitment to advancing the field. If so, join us. https://thinkingmachines.paperform.co/
Thinking Machines Lab 发表文章,主张构建延伸人类意志与判断的 AI,认为当前多数模型集中训练后固化,未受使用者塑造。公司提出训练强模型、提供可微调模型权重的工具、开发原生多模态交互模型、发布研究等四个技术方向,并将对齐视为分散在多元模型生态中的持续过程。
Mistral 发布首个具身导航模型 Robostral Navigate,8B 参数,仅靠单个普通 RGB 摄像头、无需 LiDAR 或深度传感器,在 R2R-CE validation unseen 上达到 76.6% 成功率,比最佳单摄像头方案高 9.7 分、比最佳多传感器方案高 4.5 分。
Hugging Face 宣布 vLLM 的 transformers 建模后端现在达到甚至超过 vLLM 原生实现的吞吐速度,模型作者无需移植代码即可用 --model-impl transformers 获得 vLLM 级推理性能。
推荐理由:官方给出了与 vLLM 原生实现对比的具体吞吐数字和使用方法,读者可以据此判断是否切换到自己已有的 transformers 模型工作流。
Microsoft Build 2026 上宣布 Foundry Managed Compute 以及 Hugging Face 模型精选目录,数千个开源权重模型每周更新,可一键部署到 Foundry Managed Compute,支持 NVIDIA A100、H100 和 AMD MI300X。
推荐理由:原文完整说明了精选模型目录、策展管线和运行时选择,读者可以据此评估在私有网络内部署开源模型的路径。
SkyPilot 与 Hugging Face 联合推出 store: hf,把 Hugging Face Storage 作为 SkyPilot 的一等存储后端,用一个 hf:// URL 和现有 HF_TOKEN 即可把 Bucket 或 Hub 仓库挂载进任意云上的任务。
推荐理由:两家联合发布 store: hf,给出了零出口费读取、MOUNT/COPY 用法和实测写入速度,读者可以据此评估多云 GPU 训练的存储方案。
Hugging Face 发布 LeRobot v0.6.0,引入 VLA-JEPA、FastWAM、LingBot-VA 等世界模型策略,新增 GR00T N1.7、MolmoAct2、EO-1 等 VLA,以及统一奖励模型 API(Robometer、TOPReward)。
推荐理由:发布方系统列出世界模型策略、奖励模型 API、新基准与部署 CLI 等更新,读者可以据此评估机器人学习工作流的变化。
腾讯混元 AI Infra 团队的 HPC-Ops 算子库中的 Attention 和 MoE 内核已合入 vLLM 主分支,成为一等后端,针对 NVIDIA Hopper 架构优化,在 H20 上效果最好。
Hugging Face 为 Kernels 项目引入 Hub 上的 "kernel" 新仓库类型,并默认只加载受信任发布者的内核,非信任来源需显式传入 trust_remote_code=True。
美团发布 LongCat-2.0,总参数 1.6 万亿的 MoE 大语言模型,每 token 激活约 48B,权重以 MIT 协议开源。预训练覆盖超过 35 万亿 token,训练与部署全程基于 AI ASIC 超算集群;模型引入 LongCat Sparse Attention,并在 1M 上下文数据上训练,支持 GPU 与 NPU 部署。
美团发布 LongCat-2.0,一个总参数 1.6 万亿、每 token 激活约 480 亿的 MoE 大语言模型,训练与部署完全基于 AI ASIC 超节点,预训练消耗超过 35 万亿 token。
美团 LongCat 团队发布 LongCat-2.0,总参数 1.6 万亿、每 token 激活约 480 亿的 MoE 大语言模型,权重以 MIT 协议开源。
Mistral AI 发布 Apache-2.0 协议的 Leanstral 1.5,总参数 119B、激活 6B,专注 Lean 4 形式化证明工程。
ENPIRE -> ASPIRE,我们 Physical AutoResearch 系列的第二项工作。我们正在构建机器人自我改进的组件,一次一个 /skill。
Today, we give robots a /skills library that self-evolves and compounds indefinitely! Introducing ASPIRE: a robot solving its 100th task is no longer as clueless as solving its first. Coding agents observe multimodal sensory traces from simulation and real robots, launch an evolutionary search over control programs, and distill the best know-how into an ever-expanding library. ASPIRE is a new type of continual learning: "training" is skill refinement instead of gradient descent. "Trained model" is a repo of sensorimotor skills instead of floating weights. “Distributed training” is a panel of agents each practicing a different skill instead of sharded minibatches. Here's the beauty: ASPIRE gives the tired terms "sim2real transfer" and "cross-embodiment transfer" a whole new meaning. Bridging the sim-to-real gap is notoriously brutal. An end-to-end policy has to swallow both the visual shift (sim looks toyish next to a real camera) and the subtle contact physics it never quite gets right. ASPIRE sidesteps the mess, because it doesn't ship pixels or weights across the gap, but ships the know-how. The robot still has to practice in the real world, not zero-shot, but it gets there way faster because it isn't rediscovering the strategy from scratch. Same for going single-arm to bimanual hardware, which usually requires new data and retraining from zero. ASPIRE achieves up to ~10x cut in "transfer learning” tokens (yes, tokens are the new unit of *training* compute ;) Check out our gallery of 150+ tasks and 90+ skills the robots taught themselves, all on the website! Kind of wild that we can ship the "learned weights" as an HTML page rather than a GGUF. We'll open-source the full stack so your own robot library starts compounding from ours! Deep dive in thread:
Hugging Face 上线 pi-local-router,这是一个 Pi 扩展,用于在 llama.cpp 与 Hugging Face Inference Providers 之间进行路由。该仓库以 huggingface/pi-local-router 名义发布,目前仅披露了扩展的用途,未公布参数、性能或可用性细节。
Hugging Face 上线新仓库 physics-intern-opencode-plugin,为 Opencode 中的 Physics Intern 提供安装器。正文仅说明该仓库用于在 Opencode 内安装 Physics Intern,未披露模型规模、版本号或可用方式等细节。
Hugging Face 与 Cerebras 展示了一套开源、模块化的实时语音到语音流水线:Nvidia Parakeet 做语音识别,Google DeepMind 的 Gemma 4 31B 在 Cerebras 上推理,阿里 Qwen3TTS 合成语音回复。
Together AI 宣布完成 8 亿美元 C 轮融资,投资方包括 NVIDIA、Aramco Ventures、Vista Equity、General Catalyst 等,另获超 500 MW 算力容量承诺。
推荐理由:融资公告由 CEO 亲述,同时给出客户成本案例和推理栈技术进展,可看出开源模型的实际经济性。
Google Research 发布面向表格数据分类与回归的零样本基础模型 TabFM,将表格预测框定为 ICL 问题,一次前向传播即可生成预测,无需手动训练、调参或特征工程。
We’ve been impressed with GLM-5.2 and so are introducing a $9.99/month subscription to give you 2-5x discounted access to it and other open weight models like DeepSeek, Kimi, MiniMax, Mimo, Qwen. Use it on Cline CLI & IDE with $1.99 special promo if sign up via: npm i -g cline
Every Eval Ever(EEE)与 Hugging Face Community Evals 现已互通,评测结果可跨平台发布并回溯至完整记录。EEE 数据存储已收录约 229,000 条评测结果,覆盖超 22,000 个模型和 2,200 个基准,来自 31 种报告格式。贡献者可通过转换器将 EEE 记录提交至 Community Evals,经组织官方账号提交的结果会显示验证标记。
Interconnects 第 22 期开放模型盘点指出,开源模型发布方正从少数(以中国为主的)玩家扩展为更多全球小众公司,并将发布者分为纯模型公司、大厂和产品公司三类。
Sebastian Raschka 发布长篇教程,讲解如何用 Ollama 服务 Qwen3.6 35B-A3B 等开源权重模型,并接入 Qwen Code、Codex、Claude Code 三种编码 agent harness,评测代码在 https://github.com/rasbt/local-coding-agent-evals。
vLLM-Omni 已支持 Qwen3-TTS、VoxCPM2、Fish Speech S2 Pro 和 Higgs Audio V3 等 TTS 系统,并针对各模型瓶颈采用不同优化策略。
Nathan Lambert 撰文分析 Z.ai 于 6 月 13 日向 GLM Coding Plan 用户推出、6 月 16 日以 MIT 许可发布权重的 GLM-5.2,认为它是首个在编码工具链中作为通用智能体表现称职的开源权重模型。
推荐理由:作者结合亲测与社区评测,分析 GLM-5.2 为何标志开源模型首次在编码智能体场景形成可信替代。
针对华盛顿近期签署的 AI 模型审查行政令、国会立法提案以及限制外国公民访问 Anthropic 最先进模型等监管动向,Nathan Lambert 与 Interconnected 联合撰文警告,未来政策可能误伤甚至禁止开源 AI,而这将是严重错误。
OpenAI 在 GitHub 上线新仓库 openai/planttalk,让用户借助 ChatGPT 为室内植物赋予"声音"。该仓库标题为 openai/planttalk,正文描述为"Give your houseplants a voice with ChatGPT",目前公开信息仅此一句,具体功能与用法尚未披露。