跳到正文

#推理

今日 19 条
今天10月1日周四
  1. Artificial Analysis 完整文章(网页)79

    Artificial Analysis 评测 Gemini 4 Argon:Google 重回智能前三

    Artificial Analysis 评测 Google DeepMind 新模型 Gemini 4 Argon,其在 Artificial Analysis Intelligence Index 得 53 分,追平 GPT-6 Astra(max),高于 GPT-6.1 Sol(52),为 Google 超 7 个月来首个高于 Flash 档的专有模型。

    推荐理由:第三方评测给出了智能指数、单位任务成本、幻觉率等多项横向数据,可用于比较 Gemini 4 Argon 与竞品的实际表现。

  2. Rohan Paul44

    美国 Austin 的 webAI 推出 3.66B 本地模型 TwIL-LM3-Pro,将 IBM Granite 的形式逻辑分数提升 28%,在形式逻辑上比 VibeThinker-3B 高约 35%、比 Qwen3.5-4B 高 24%、比 LFM2.5-8B-A1B 高 47%。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  3. Google DeepMind:Blog(RSS)77

    Google DeepMind 发布 Gemini 4 Argon 前沿模型

    Google DeepMind 宣布新前沿模型 Gemini 4 Argon,先通过 Fairwind Program 向可信网络防御者开放,再逐步扩展至开发者、企业和消费者。

    推荐理由:原文给出定价、1M 输出上限和多项基准成绩,读者可据此评估该模型在编码与防御性网络安全上的实际表现。

  4. Google Blog:AI(RSS)76

    Google 发布 Gemini 4 Argon 前沿模型

    Google 发布新前沿模型 Gemini 4 Argon,先通过 Fairwind Program 面向可信网络防御者开放,价格为每百万输入 token $2、输出 token $10,缓存输入 token 为输入价的 5%。

    推荐理由:官方公告给出定价、输出 token 上限和多个基准分数,读者可以据此评估它在编码与安全防御场景的落点。

  5. Rohan Paul63

    Google 发布 Gemini 4 Argon,Sundar Pichai 称其在复杂工作流、网络防御和软件工程上表现前沿。

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  6. Rohan Paul61

    Google 发布新旗舰模型 Gemini 4 Argon,作者称其在多数基准上超过 GPT-6 Astra 和 Claude Opus 5.5,输出上限从 64K 提升到行业领先的 1M tokens,约为此前 128K 上限的近 8 倍。引用材料提到其在 Harvey's Legal Agent Benchmark 上领先,且目前仅限 Google 员工、经审核的网络安全相关机构和可信测试者使用。

    引用Rohan Paul@rohanpaul_ai

    MASSIVE reveal from Google. Its new flagship, Gemini 4 Argon, outscores GPT-6 Astra and Claude Opus 5.5 on most benchmarks. - beats GPT-6 Astra and Claude Opus 5.5 on some super important industry benchmarks. - its widest lead in legal work, 19.6% on Harvey's Legal Agent Benchmark against 6.7% for Anthropic's Claude Fable 5.1. - output limit jumps from 64K to 1M tokens, an industry-leading ceiling, - Only 3 groups have it today. the first is Google's own staff, vetted cyber defenders such as government agencies and security companies and trusted testers giving Google feedback. - Inside Google, Argon agents freed over 300 TiB of data-center memory, with 500 TiB to 1 PiB of total savings estimated, and made a Rust port of the libgav1 video decoder 2.7x faster by replacing 32K lines of SIMD code.

  7. Arena.ai66

    Google DeepMind 发布新前沿模型 Gemini 4 Argon,通过 Fairwind Program 向部分受信任测试者开放。

    引用Google DeepMind@GoogleDeepMind

    Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.

    推荐理由:榜单方公布了 Gemini 4 Argon (High) 在 Text Arena 的分项名次、1525 分和混合价格,读者可据此对比成本效率。

  8. 🚨 AI News | TestingCatalog47

    突发 🔥:Google 宣布 Gemini 4 Argon,一款新的前沿模型,面向"跨真实世界软件工程、法律和金融等企业知识工作、以及网络防御的复杂工作流"。 在 DeepSWE v1.1 上取得 77.9% 的分数,创下新 SOTA。在众多基准测试上表现优于 GPT-6 Astra、Opus 5.5 和 Fable 5.1。 即将推出,首先面向付费 API 客户和 Google AI Ultra 订阅用户。 很快!👀

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

9月30日周三
  1. Prime Intellect(网页)66

    Prime Intellect 发布 INTELLECT-2:首个通过全球分布式强化学习训练的 32B 模型

    Prime Intellect 发布 INTELLECT-2,称其为首个通过全球分布式强化学习训练的 32B 参数推理模型,在异步、异构的分布式算力网络上训练,性能较 QwQ-32B 在数学与编码基准上有提升。

    推荐理由:官方发布首个通过全球分布式强化学习训练的 32B 推理模型,并开源训练框架、代码和数据,适合关注开源分布式训练路线的读者。

  2. Prime Intellect(网页)71

    Prime Intellect 发布 106B MoE 模型 INTELLECT-3 并开源完整训练配方

    Prime Intellect 发布 INTELLECT-3,一个 106B 参数的 MoE 模型,在 GLM 4.5 Air 基座上经过 SFT 和大规模 RL 训练,在数学、代码、科学和推理基准上达到同规模最先进水平,超越许多更大的前沿模型,如 AIME24 90.8 分、AIME25 88.0 分、GPQA-Diamond 74.4 分。

    推荐理由:原文同时放出模型权重、训练框架和全部环境,读者可以看到大规模 RL 后训练的完整可复用配方。

  3. MiniMax:Blog(网页)75

    MiniMax 发布 M2.5 模型,SWE-Bench Verified 达 80.2%

    MiniMax 发布 MiniMax-M2.5,在 SWE-Bench Verified 得分 80.2%、Multi-SWE-Bench 51.3%、BrowseComp 76.3%,完成 SWE-Bench Verified 评测比 M2.1 快 37%,与 Claude Opus 4.6 的 22.9 分钟持平。

    推荐理由:官方公布了 M2.5 在编码、搜索和办公场景的成绩与价格细节,读者可以据此比较它与同类模型的性价比。

  4. LMSYS:Blog(Chatbot Arena 团队)72

    SGLang 和 Miles 实现 NVIDIA Nemotron 3 Ultra Day-0 支持

    SGLang 和 Miles 宣布对 NVIDIA Nemotron 3 Ultra 提供 Day-0 支持。该模型为开源 MoE 混合 Transformer-Mamba 架构,总参数 550B、激活 55B,上下文最长 1M tokens,支持 NVFP4 与 BF16 权重,面向长程自主智能体优化。

    推荐理由:原文给出架构、部署配置和 RL 训练细节,读者可以据此评估该模型在长程智能体场景的可用性。

  5. Cursor Blog69

    Cursor 与 SpaceXAI 联合发布 Grok 4.6,聚焦长时程智能体任务

    Cursor 团队与 SpaceXAI 今日联合发布 Grok 4.6,基于 Grok 4.5 打造,专注于长时间运行的智能体、更复杂的交互与视觉工作。在 Artificial Analysis Intelligence Index 上它与 GPT-5.6 Sol 持平(61 分),训练上进行了更长时间的补充训练、重新生成的 SFT 轨迹和广泛的智能体 RL。

    推荐理由:原文给出训练方法、基准分数和定价细节,读者可以据此评估它在智能体编程场景中的定位。

  6. Cognition 模型 / Devin 博客(网页)65

    Cognition 发布 SWE-grep 和 SWE-grep-mini,用 RL 训练快速并行上下文检索模型

    Cognition 训练了 SWE-grep 和 SWE-grep-mini 两个快速智能体模型,专攻高度并行的上下文检索,检索能力匹敌前沿编码模型,耗时低一个数量级。

    推荐理由:官方公布了模型设计、RL 训练细节和评测数字,读者可以借此了解并行上下文检索的实现路径。

  7. Artificial Analysis 完整文章(网页)75

    Artificial Analysis 评测 GPT-6 Astra:与 Claude Fable 5.1 并列两大指数第一且成本更低

    Artificial Analysis 发布 GPT-6 Astra 基准测试报告,该模型在 Intelligence Index 得 53 分、Coding Agent Index 得 62 分,均与 Claude Fable 5.1 并列第一,且成本分别约为其 40% 和 60%。

    推荐理由:原文给出 GPT-6 Astra 在两大指数中的得分、成本和 token 效率数据,可据此比较它与 Claude Fable 5.1 的实际表现。

  8. ARC Prize:官方博客68

    ARC Prize 测评 OpenAI o1 在 ARC-AGI-Pub 的表现

    ARC Prize 用同一基线测试框架实测 OpenAI o1-preview 和 o1-mini:o1-preview 在 ARC-AGI 公开评测得 21.2%,与 Claude 3.5 Sonnet(21%)相当,但单任务平均耗时 4.2 分钟,约为 Sonnet 的 10 倍,400 道公开题共花 70 小时,而 GPT-4o 和 Sonnet 只需 30 分钟。

    推荐理由:ARC Prize 用同一测试框架给出 o1 在 ARC-AGI-Pub 的基线分数和耗时对比,并分析了测试时计算扩展的效率问题。

  9. ARC Prize:官方博客69

    ARC Prize 实测 o3 与 o4-mini 在 ARC-AGI 上的表现

    ARC Prize Foundation 发布 o3 和 o4-mini 在 ARC-AGI 上的首次公开评测:o3-medium 在 ARC-AGI-1 Semi Private Eval 得 53%,o4-mini-medium 得 42%,两者在 ARC-AGI-2 上均低于 3%。

    推荐理由:ARC Prize 官方实测给出了 o3 与 o4-mini 在 ARC-AGI 两代基准上的准确率、成本和 high 推理不完整覆盖的数据细节。