跳到正文

全部动态

今日 46 条
9月9日周三
  1. Mark Chen60

    OpenAI 宣布给出 Navier-Stokes 千禧年大奖难题的一个证明,该问题关于三维光滑流体运动的描述是否会失稳,已悬置约 90 年。证明由一组智能体使用一个能力显著强于 GPT-6 Astra 的 OpenAI 下一代模型产出。OpenAI 首席研究官 Mark Chen 转发并表示,许多同事因相信 AI 是解决各领域大挑战的最快途径而离开原领域,看到第一个挑战落下令他感到不真实。

    引用OpenAI@OpenAI

    We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.

9月8日周二
  1. Google DeepMind:Blog(RSS)80

    Google DeepMind 发布 AlphaGenome Atlas,覆盖人类基因组 90 亿个单核苷酸变异预测

    Google DeepMind 发布 AlphaGenome Atlas,一个包含人类基因组全部约 90 亿个单核苷酸变异效应预测的平台,规模达 1PB,是 AlphaFold Database 的 30 倍以上。

    推荐理由:AlphaGenome Atlas 把 90 亿个单核苷酸变异的预测结果做成可检索资源,读者可了解其数据规模与在罕见病研究中的验证案例。

9月7日周一
9月6日周日
9月5日周六
  1. AI at Meta54

    Meta 宣布 Muse Spark 1.3 的 max 推理档现已上线 Muse Code 和 Meta Model API。官方称该档在编码和智能体任务上表现明显更强,建议即使已试过 1.3 high 或 1.3 xhigh 也再试一次。

    引用Alexandr Wang@alexandr_wang

    1/ we just publicly released Muse Spark 1.3 max! we see significantly stronger coding and agentic performance on muse spark 1.3 max, so would strongly recommend trying it out even if you've already tried muse spark 1.3 high or muse spark 1.3 xhigh.

9月4日周五
  1. Unsloth AI71

    Unsloth 发布 GLM-5.3-Flash 的本地 GGUF 优化方案,通过更快解码和 MTP 支持使本地推理提速 1.6 至 3.4 倍,长上下文下最高达 3.3 倍。3-bit 量化可在 128GB 内存设备上通过 Unsloth Desktop 或 llama.cpp 运行,GGUF 权重和指南已发布,开箱即用且无需额外模块或 MTP 文件。

    引用Z.ai@Zai_org

    Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: http://z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: http://huggingface.co/zai-org/GLM-5.3-Flash API: http://docs.z.ai/guides/llm/glm-5.3-flash Coding Plan: http://z.ai/subscribe ZCode: http://zcode.z.ai/en Chat: http://chat.z.ai AutoClaw: http://autoclaw.z.ai

    推荐理由:原文给出本地推理提速的具体配置和硬件门槛,读者可以直接复用到自己的 GLM-5.3-Flash 部署中。

  2. Mark Chen80

    OpenAI 首席研究官 Mark Chen 宣布 GPT-6 Astra 发布,称其汇集多年预训练、强化学习和后训练工作,是该团队迄今能力最强、对齐程度最高的模型。

    引用OpenAI@OpenAI

    This is GPT-6 Astra. Anything you can do on a computer, Astra can do for you. Fast.

    推荐理由:OpenAI 首席研究官亲自说明 GPT-6 Astra 的能力变化与对齐工作,可帮助读者了解官方对 Computer Use 和 Agent 监督的进展表述。

9月3日周四
  1. Google DeepMind:Blog(RSS)74

    Google DeepMind 发布 WeatherNext 3 全球天气 AI 模型

    Google DeepMind 与 Google Research 发布 WeatherNext 3 全球天气 AI 模型,直接学习实时地球静止卫星数据,每小时生成一次预报,地表变量分辨率达 5 公里,整体比 WeatherNext 2 的 25 公里网格清晰约 5 倍。

    推荐理由:对比前代的分辨率、更新频率与降水评分提升,可了解实时卫星数据如何改变全球天气预报的精度边界。

  2. Demis Hassabis59

    Demis Hassabis 宣布发布 Gemini 3.8 Flash,这是不到一个月内的又一次升级,主要提升智能体与编码能力;同时推出面向网络防御前沿的 3.8 Flash Cyber。模型详情和安全工作见博客:https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/。

    引用Logan Kilpatrick@OfficialLoganK

    Introducing Gemini 3.8 Flash, another jump in Gemini's agentic + coding capabilities, and our 3rd updated Flash model in only 6 weeks... This model has been a ton of fun to work with, excited to see what you all think!

9月2日周三
  1. Hao AI Lab46

    很高兴看到 @Physion_Labs 对 FastH3 Preview、@MiniMax_AI 和 H3 Max 的独立评测 > FastH3 preview 作为社区开源成果,整体表现相当不错:它保持了很强的提示词遵循度,在这一维度上甚至超过了 H3…… 敬请期待我们下一版 FastH3,质量会更高!

    引用Physion Labs Official@Physion_Labs

    🐶🏁 Three models. One race. We independently tested H3 Max by @fal, minimax H3 by @MiniMax_AI, and FastH3 preview by @haoailab @haozhangml @wlsaidhi across robotics, animation, movies, and ads. H3 Max takes the overall lead 🏆. FastH3_preview, meanwhile, is a community OSS effort that’s already keeping up surprisingly well, even beating H3 on overall Prompt Adherence 🥳 The bigger gaps show up in Visual Integrity and Human Preference. Full evaluation: https://physionlabs.ai/blog/minimax-h3-evaluation

  2. Fei-Fei Li57

    Fei-Fei Li 宣布 World Labs 团队发布多模态世界模型 Atlas,称其从零训练、可生成像素级精准相机控制帧。Atlas 还能从单张输入图像重建大场景、通过重排视频帧模拟时空、从一张或多张图像原生输出 3D 空间,并将多张带位姿图像合成一致的 3D 世界,应用方向覆盖 VFX 到机器人。

    引用World Labs@theworldlabs

    Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.

9月1日周二
8月30日周日
8月29日周六
8月28日周五
  1. SiliconFlow66

    智谱(Z.ai)的 GLM-5.3 已开源,并上线 SiliconFlow,提供 Day-0 支持。该模型总参数 744B、激活 40B,与 GLM-5.2 同底座,经后训练大幅提升,官方称在编程与智能体基准上可与 Fable 5 和 GPT-5.6 Sol 竞争。SiliconFlow 还提到其在 OpenRouter 上 GLM-5.2 token 份额排名第一。

    推荐理由:原文给出了 GLM-5.3 的参数规模、开源状态和基准表现,读者可以据此评估是否迁移现有工作流。

8月27日周四
  1. NVIDIA Technical Blog(开发者技术博客 · RSS)65

    阿里发布 Qwen3.8-Flash-Next 权重,预览 Qwen4 架构

    阿里巴巴发布 Qwen3.8-Flash-Next 模型权重,作为即将到来的 Qwen4 架构的预览,供开发者实验和评估。该模型是多模态混合专家(MoE)模型,主模型 125B 参数,附加 51B N-gram embeddings,每 token 激活 6B 参数;原生上下文窗口 262,144 token,可扩展至 1M token。

8月26日周三
8月25日周二
  1. Hugging Face:Blog(RSS)66

    IBM 发布 Granite 4.2 推理模型系列并详解构建过程

    IBM Granite 团队发布 Granite 4.2 推理模型系列,包含 3B、8B、30B 三个 dense 版本,基于 Granite-4.1 基座(约 15T tokens 预训练,上下文扩展到 512K),经 SFT 和多阶段 GRPO 强化学习训练,8B 和 30B 额外经历 SWE、终端、搜索三类真实环境 agentic RL。

    推荐理由:官方完整披露了从预训练、SFT 到多阶段 GRPO 强化学习的训练细节,读者可以据此了解推理模型的完整构建流程。