跳到正文

全部动态

今日 395 条
今天10月1日周四
  1. Arena.ai78

    Arena 公布 Gemini 4 Argon (High) 在 Agent Arena 以 +7.92% 净改进分排名第 8,每任务成本 $0.62,重塑了 Pareto 前沿,比 Gemini 3.8 Flash (High) 高 4.96 个百分点。

    引用Arena.ai@arena

    Big news: Gemini 4 Argon (High) by @GoogleDeepMind just landed #1 in Text Arena with 1525 pts, and #8 in Code Arena: WebDev with 1679 pts! This release has reshaped the Text Arena Pareto frontier with a blended $8/MToken! Gemini 4 Argon (High) is now the most cost efficient model, see its placement on Pareto frontier below. In the Text Arena, Gemini 4 Argon (High) ranks #1 in Coding, Hard Prompts, Instruction Following, Longer Query, and Creative Writing. It also leads every occupational domain evaluated, with additional #1 spots in English, Non-English, Chinese, and Russian. This model is +20 points above the #2 ranked Claude Opus 4.6 (High), and a huge leap from Google’s previous release, Gemini 3.8 Flash (High) at #11! In Code Arena: WebDev, Gemini 4 Argon (High) gained +96 points from Gemini 3.8 Flash (High), and went from #29 to #8. Congrats to the @GoogleDeepMind team on this impressive frontier release!

    推荐理由:原文给出 Agent Arena 排名、关键信号得分和每任务成本数据,读者可以据此评估该模型在真实智能体任务中的性价比。

  2. Every:最新文章(网页)36

    Sam Altman 如何用 OpenAI 的 Dots 智能体夺回时间

    OpenAI CEO Sam Altman 在 DevDay 后接受 The Every Podcast 采访,讲述他如何用 OpenAI 新的常驻智能体 Dot 安排日程、节省时间,并称自己离不开 Astra 的 Ultrafast 模式。本届 DevDay 共发布 22 项产品与功能,数量是去年的两倍多,Altman 还谈到自己如何构建新功能,以及为何相信 AI 将带来新的文艺复兴。

  3. Rohan Paul44

    美国 Austin 的 webAI 推出 3.66B 本地模型 TwIL-LM3-Pro,将 IBM Granite 的形式逻辑分数提升 28%,在形式逻辑上比 VibeThinker-3B 高约 35%、比 Qwen3.5-4B 高 24%、比 LFM2.5-8B-A1B 高 47%。

    引用David Stout@Davidstout

    Half a million downloads in a month. Today, our open source family takes another step forward. Thank you for the incredible support behind our first-generation models. We’re excited to introduce TwIL-LM3-Pro. At just 3.6 billion parameters, it brings powerful reasoning to everyday computers, with quantized builds that run locally. No cloud required. In our evaluation: Formal logic: Highest recorded headline score among the small models compared—beating China’s VibeThinker-3B by 35% and Qwen3.5-4B by 24%, and Liquid AI’s LFM2.5-8B-A1B by 47%. Broader reasoning: 95% on SVAMP and 64.1% on MuSR, the highest recorded scores among the small models compared. BIG-Bench Hard’s logic subset: 95.4%, compared with VibeThinker-3B’s 61.1%. We believe AI is entering a post-training era. The advantage will increasingly belong to companies with the best pipelines and those that can produce capable, personalized intelligence faster and more efficiently, then put it on devices people already own. That’s what we’re building at webAI. And we’re only beginning to share what’s coming out of our lab. Coming soon: Meridian, our family of frontier-class models built to run on device. Our most advanced models will be available through the @thewebAI application. Join the waitlist as we expand access. Proudly built in Austin, Texas. 🇺🇸

  4. Bloomberg:Technology(RSS)58

    美光因 AI 内存需求旺盛发布超预期季度业绩指引

    美光(Micron Technology)发布本财季指引,预计截至 11 月的财季营收约 615 亿美元,高于分析师平均预期的 568 亿美元;剔除部分项目后每股利润约 38.15 美元,高于预期的 36.02 美元。公司称 AI 建设热潮推动内存需求空前旺盛。

  5. Bloomberg:Technology(RSS)41

    Micron 业绩指引超预期:关键要点解读

    Micron 给出的本季度销售指引大幅超出市场预期,AI 建设热潮推动内存需求达到前所未有的水平。Gabelli 分析师 Hendi Susanto 指出,投资者要求 Micron 下一财年营收和利润翻倍,超大规模云厂商正提前锁定多年内存供应并附带定价窗口,他认为本轮周期与以往内存行业繁荣萧条有所不同。

  6. Bloomberg:Technology(RSS)46

    HPE 因网络业务前景向好上调业绩预期,获 Vultr 12 亿美元订单

    慧与(HPE)上调网络业务营收预期,并宣布获得云公司 Vultr 价值 12 亿美元的订单,显示其正受益于 AI 基础设施建设。HPE 去年收购 Juniper Networks 后扩展了网络业务,现预计该部门 2027 财年销售额增长将在高双位数至低 20% 区间,长期销售展望为高双位数,运营利润率目标最高达高 20% 区间。

  7. Google DeepMind:Blog(RSS)77

    Google DeepMind 发布 Gemini 4 Argon 前沿模型

    Google DeepMind 宣布新前沿模型 Gemini 4 Argon,先通过 Fairwind Program 向可信网络防御者开放,再逐步扩展至开发者、企业和消费者。

    推荐理由:原文给出定价、1M 输出上限和多项基准成绩,读者可据此评估该模型在编码与防御性网络安全上的实际表现。

  8. Google AI:DEV 作者专属(RSS)48

    用三对象 CRM 数据模型终结数据孤岛

    收入运营顾问 Alex 提出用 Company、Contact、Deal 三个一等对象重构 CRM 数据模型,其余信息一律降为属性或事件,以消除营销、销售、客户成功三套心智模型写入同一数据库造成的孤岛。该模型让关联查询变成两跳、状态机清晰可读、自动化不再冲突,并让 AI 智能体有明确状态可推理。迁移采用 strangler-fig 方式,新模型与旧字段并行运行,旧属性隐藏一个季度后再删除。

  9. Google AI:DEV 作者专属(RSS)23

    One Commit a Day 六个月的 OSS 回顾:发布 v2.0 与 v2.0.1

    开发者发布 One Commit a Day v2.0 和 v2.0.1,重构仓库并扩充日志与模板,同时完成 Doc MCP Server 的重构与重构,并处理了 Dependabot 与 MCP SDK 维护。该项目已坚持"每天一次提交"六个月,作者称一致性不等于每天同等强度,九月有不少"恢复模式"的日子。未来六个月他计划加入一个更大、更活跃的开源项目并推动自有项目走向生产可用。

  10. Google AI:DEV 作者专属(RSS)37

    OpenAI 将 AI 工作研究落地:携手 America's SBDC 为小企业团队提供实操培训

    OpenAI 宣布与 America's SBDC 达成全国性合作试点,为 SBDC 顾问提供 OpenAI Academy 培训,并在各州开展线下工作坊,预计超过 1000 家小企业参与,参与者可免费使用 ChatGPT Plus。该计划配合其 Work at the Frontier 报告,聚焦客户互动、产品开发和财务等可由 AI 智能体承接的实际工作场景。

  11. Google AI:DEV 作者专属(RSS)46

    读者指出我的修复并未解决智能体等待问题

    一位读者纠正了作者此前提出的单行修复方案:把 CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS 设为 0 只是取消截止时间,而非让等待变得可追踪。读者建议每个延迟任务都应留下机器持有的记录,包含任务 id、明确截止时间和到期后的下一步动作,并由独立机制核对。作者已将其转为新项目模板中的 ticket,但该 ticket 已挂起七天,修复尚未落地。

  12. Google AI:DEV 作者专属(RSS)52

    Verax 的 Agent 权限策略:没有规则时默认拒绝

    Verax 对 AI Agent 的请求采取默认拒绝策略,没有规则的调用一律拒绝,包括 Agent 换工具名重试的情况。策略只列 memory.get、memory.put、audit.explain、message.read 四个工具,同一工具出现两条规则会在加载时被拒绝;拒绝记录与批准记录同样签名留档,可用 verax verify 离线验证。

  13. Google Blog:AI(RSS)76

    Google 发布 Gemini 4 Argon 前沿模型

    Google 发布新前沿模型 Gemini 4 Argon,先通过 Fairwind Program 面向可信网络防御者开放,价格为每百万输入 token $2、输出 token $10,缓存输入 token 为输入价的 5%。

    推荐理由:官方公告给出定价、输出 token 上限和多个基准分数,读者可以据此评估它在编码与安全防御场景的落点。

  14. Karina50

    Gemini 的 PostTrainBench 得分翻了一倍多:21.99%(3.1 Pro)→ 45.3%(4)🔥🚀

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  15. 🚨 AI News | TestingCatalog56

    Inworld 宣布收购 Ultravox,一个用于构建实时语音智能体的平台,可听、推理、调用工具并在对话中处理轮次和打断。Ultravox 上内置的 Inworld 声音现已运行在 Realtime TTS-2 上,现有 voice ID 无需迁移或修改代码即可继续使用,TTS-2 还为 Ultravox 智能体增加自然语言引导能力。

    引用Inworld AI@inworld

    We’re excited to announce that @ultravox_dot_ai is now part of Inworld. Ultravox is the platform developers use to build real-time voice agents. We've worked with the team for a while through our TTS partnership, and today members of the team that built it are joining Inworld to keep developing it.

  16. Rohan Paul63

    Google 发布 Gemini 4 Argon,Sundar Pichai 称其在复杂工作流、网络防御和软件工程上表现前沿。

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  17. Rohan Paul61

    Google 发布新旗舰模型 Gemini 4 Argon,作者称其在多数基准上超过 GPT-6 Astra 和 Claude Opus 5.5,输出上限从 64K 提升到行业领先的 1M tokens,约为此前 128K 上限的近 8 倍。引用材料提到其在 Harvey's Legal Agent Benchmark 上领先,且目前仅限 Google 员工、经审核的网络安全相关机构和可信测试者使用。

    引用Rohan Paul@rohanpaul_ai

    MASSIVE reveal from Google. Its new flagship, Gemini 4 Argon, outscores GPT-6 Astra and Claude Opus 5.5 on most benchmarks. - beats GPT-6 Astra and Claude Opus 5.5 on some super important industry benchmarks. - its widest lead in legal work, 19.6% on Harvey's Legal Agent Benchmark against 6.7% for Anthropic's Claude Fable 5.1. - output limit jumps from 64K to 1M tokens, an industry-leading ceiling, - Only 3 groups have it today. the first is Google's own staff, vetted cyber defenders such as government agencies and security companies and trusted testers giving Google feedback. - Inside Google, Argon agents freed over 300 TiB of data-center memory, with 500 TiB to 1 PiB of total savings estimated, and made a Rust port of the libgav1 video decoder 2.7x faster by replacing 32K lines of SIMD code.