跳到正文

#推理

今日 25 条
今天10月1日周四
  1. HuggingFace Daily Papers(社区热门论文)39

    ThinkV2V:释放 MLLM 推理能力,实现指令引导的视频编辑

    ThinkV2V 是一个推理驱动的指令引导视频编辑框架,在视觉生成前显式激活 MLLM 思考,通过 MLLM-to-DiT 架构将思考转化为精炼的条件信号。它结合渐进式课程训练与推理时思考扩展,并构建了 ThinkV2V-150K 数据集和 ThinkV2V-Bench 基准。实验显示其 5B 规模 DiT 模型在复杂与标准编辑场景均达 SOTA,显著超越更大的 10B 规模基线。

  2. Artificial Analysis 完整文章(网页)79

    Artificial Analysis 评测 Gemini 4 Argon:Google 重回智能前三

    Artificial Analysis 评测 Google DeepMind 新模型 Gemini 4 Argon,其在 Artificial Analysis Intelligence Index 得 53 分,追平 GPT-6 Astra(max),高于 GPT-6.1 Sol(52),为 Google 超 7 个月来首个高于 Flash 档的专有模型。

    推荐理由:第三方评测给出了智能指数、单位任务成本、幻觉率等多项横向数据,可用于比较 Gemini 4 Argon 与竞品的实际表现。

  3. Rohan Paul44

    美国 Austin 的 webAI 推出 3.66B 本地模型 TwIL-LM3-Pro,将 IBM Granite 的形式逻辑分数提升 28%,在形式逻辑上比 VibeThinker-3B 高约 35%、比 Qwen3.5-4B 高 24%、比 LFM2.5-8B-A1B 高 47%。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  4. Google DeepMind:Blog(RSS)77

    Google DeepMind 发布 Gemini 4 Argon 前沿模型

    Google DeepMind 宣布新前沿模型 Gemini 4 Argon,先通过 Fairwind Program 向可信网络防御者开放,再逐步扩展至开发者、企业和消费者。

    推荐理由:原文给出定价、1M 输出上限和多项基准成绩,读者可据此评估该模型在编码与防御性网络安全上的实际表现。

  5. Google Blog:AI(RSS)76

    Google 发布 Gemini 4 Argon 前沿模型

    Google 发布新前沿模型 Gemini 4 Argon,先通过 Fairwind Program 面向可信网络防御者开放,价格为每百万输入 token $2、输出 token $10,缓存输入 token 为输入价的 5%。

    推荐理由:官方公告给出定价、输出 token 上限和多个基准分数,读者可以据此评估它在编码与安全防御场景的落点。

  6. Rohan Paul63

    Google 发布 Gemini 4 Argon,Sundar Pichai 称其在复杂工作流、网络防御和软件工程上表现前沿。

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  7. Rohan Paul61

    Google 发布新旗舰模型 Gemini 4 Argon,作者称其在多数基准上超过 GPT-6 Astra 和 Claude Opus 5.5,输出上限从 64K 提升到行业领先的 1M tokens,约为此前 128K 上限的近 8 倍。引用材料提到其在 Harvey's Legal Agent Benchmark 上领先,且目前仅限 Google 员工、经审核的网络安全相关机构和可信测试者使用。

    引用Rohan Paul@rohanpaul_ai

    MASSIVE reveal from Google. Its new flagship, Gemini 4 Argon, outscores GPT-6 Astra and Claude Opus 5.5 on most benchmarks. - beats GPT-6 Astra and Claude Opus 5.5 on some super important industry benchmarks. - its widest lead in legal work, 19.6% on Harvey's Legal Agent Benchmark against 6.7% for Anthropic's Claude Fable 5.1. - output limit jumps from 64K to 1M tokens, an industry-leading ceiling, - Only 3 groups have it today. the first is Google's own staff, vetted cyber defenders such as government agencies and security companies and trusted testers giving Google feedback. - Inside Google, Argon agents freed over 300 TiB of data-center memory, with 500 TiB to 1 PiB of total savings estimated, and made a Rust port of the libgav1 video decoder 2.7x faster by replacing 32K lines of SIMD code.

  8. Arena.ai66

    Google DeepMind 发布新前沿模型 Gemini 4 Argon,通过 Fairwind Program 向部分受信任测试者开放。

    引用Google DeepMind@GoogleDeepMind

    Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.

    推荐理由:榜单方公布了 Gemini 4 Argon (High) 在 Text Arena 的分项名次、1525 分和混合价格,读者可据此对比成本效率。

  9. 🚨 AI News | TestingCatalog47

    突发 🔥:Google 宣布 Gemini 4 Argon,一款新的前沿模型,面向"跨真实世界软件工程、法律和金融等企业知识工作、以及网络防御的复杂工作流"。 在 DeepSWE v1.1 上取得 77.9% 的分数,创下新 SOTA。在众多基准测试上表现优于 GPT-6 Astra、Opus 5.5 和 Fable 5.1。 即将推出,首先面向付费 API 客户和 Google AI Ultra 订阅用户。 很快!👀

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  10. HuggingFace Daily Papers(社区热门论文)56

    研究揭示分块 KV-cache 压缩导致模型出现周期性相位敏感弱位

    论文发现采用分块 KV-cache 压缩的模型存在相位敏感,检索准确率随压缩 stride 周期性波动,弱位置足以翻转答案。在 DeepSeek-V4 系列上,128K tokens 下最优与最差相位组差距达 15 到 40 个点;从零预训练的 0.6B 对照实验显示周期始终跟随 stride,full attention 各位置差距仅 6.1 个点而压缩模型可达 78 个点。

9月30日周三
  1. Berkeley RDI:Blog(AI 安全与评测)51

    Berkeley RDI 提出 DELTA 与 RL Grokking Recipe:RL 可让 LLM 解出此前完全不会的任务

    UC Berkeley 等机构发布论文(arXiv:2509.21016),提出合成编程任务套件 DELTA 与两阶段奖励策略,证明 RL 可以让基础模型在 pass@128=0 的任务上出现 grokking 式相变,准确率跳升至约 100%。方法先用逐测试密集奖励脱离全零区,再切换二值全通过奖励巩固策略;迁移实验显示学会的技能可组合外推,但在需要新不变量的变换性变化上较弱。

  2. Prime Intellect(网页)66

    Prime Intellect 发布 INTELLECT-2:首个通过全球分布式强化学习训练的 32B 模型

    Prime Intellect 发布 INTELLECT-2,称其为首个通过全球分布式强化学习训练的 32B 参数推理模型,在异步、异构的分布式算力网络上训练,性能较 QwQ-32B 在数学与编码基准上有提升。

    推荐理由:官方发布首个通过全球分布式强化学习训练的 32B 推理模型,并开源训练框架、代码和数据,适合关注开源分布式训练路线的读者。

  3. Prime Intellect(网页)71

    Prime Intellect 发布 106B MoE 模型 INTELLECT-3 并开源完整训练配方

    Prime Intellect 发布 INTELLECT-3,一个 106B 参数的 MoE 模型,在 GLM 4.5 Air 基座上经过 SFT 和大规模 RL 训练,在数学、代码、科学和推理基准上达到同规模最先进水平,超越许多更大的前沿模型,如 AIME24 90.8 分、AIME25 88.0 分、GPQA-Diamond 74.4 分。

    推荐理由:原文同时放出模型权重、训练框架和全部环境,读者可以看到大规模 RL 后训练的完整可复用配方。