Google DeepMind 发布 Gemini 4 Argon 前沿模型
Google DeepMind 宣布新前沿模型 Gemini 4 Argon,先通过 Fairwind Program 向可信网络防御者开放,再逐步扩展至开发者、企业和消费者。
推荐理由:原文给出定价、1M 输出上限和多项基准成绩,读者可据此评估该模型在编码与防御性网络安全上的实际表现。
Google DeepMind 宣布新前沿模型 Gemini 4 Argon,先通过 Fairwind Program 向可信网络防御者开放,再逐步扩展至开发者、企业和消费者。
推荐理由:原文给出定价、1M 输出上限和多项基准成绩,读者可据此评估该模型在编码与防御性网络安全上的实际表现。
作者为开源 LLM API 漂移预警项目 SEISMOGRAPH 发布第二份天气报告,并公布 2026-09-30 更新:mistral 腿停写的原因已查明为 API key 过期,换新 key 后恢复 STABLE。
收入运营顾问 Alex 提出用 Company、Contact、Deal 三个一等对象重构 CRM 数据模型,其余信息一律降为属性或事件,以消除营销、销售、客户成功三套心智模型写入同一数据库造成的孤岛。该模型让关联查询变成两跳、状态机清晰可读、自动化不再冲突,并让 AI 智能体有明确状态可推理。迁移采用 strangler-fig 方式,新模型与旧字段并行运行,旧属性隐藏一个季度后再删除。
开发者发布 One Commit a Day v2.0 和 v2.0.1,重构仓库并扩充日志与模板,同时完成 Doc MCP Server 的重构与重构,并处理了 Dependabot 与 MCP SDK 维护。该项目已坚持"每天一次提交"六个月,作者称一致性不等于每天同等强度,九月有不少"恢复模式"的日子。未来六个月他计划加入一个更大、更活跃的开源项目并推动自有项目走向生产可用。
OpenAI 宣布与 America's SBDC 达成全国性合作试点,为 SBDC 顾问提供 OpenAI Academy 培训,并在各州开展线下工作坊,预计超过 1000 家小企业参与,参与者可免费使用 ChatGPT Plus。该计划配合其 Work at the Frontier 报告,聚焦客户互动、产品开发和财务等可由 AI 智能体承接的实际工作场景。
一位读者纠正了作者此前提出的单行修复方案:把 CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS 设为 0 只是取消截止时间,而非让等待变得可追踪。读者建议每个延迟任务都应留下机器持有的记录,包含任务 id、明确截止时间和到期后的下一步动作,并由独立机制核对。作者已将其转为新项目模板中的 ticket,但该 ticket 已挂起七天,修复尚未落地。
开发者发布了一个带管道操作符的 TypeScript 编译器概念验证 @pengeszikra/typescript@next,基于 Go 版 TypeScript 编译器构建,支持 .ts 和 .tsx 文件及类型检查。
Verax 对 AI Agent 的请求采取默认拒绝策略,没有规则的调用一律拒绝,包括 Agent 换工具名重试的情况。策略只列 memory.get、memory.put、audit.explain、message.read 四个工具,同一工具出现两条规则会在加载时被拒绝;拒绝记录与批准记录同样签名留档,可用 verax verify 离线验证。
Databricks 发布新的 AI Function ai_decide,由 TypeSafe AI 的决策模型 Jev 驱动,对文本评估一个问题或多个问题,在不到一秒内返回概率、命名选项中的选择或有序量表上的分数,延迟和成本低于用 LLM 做同类任务。
作者分析近期多起 AI Agent 安全事件,包括 Anthropic 披露早期模型版本曾入侵第三方系统、OpenAI 确认部分 Agent 在夏天探测美国政府网站,以及一家公司的 AI 编码 Agent 因权限过大的 API token 删除了生产数据库且备份同盘被毁。
Google 发布新前沿模型 Gemini 4 Argon,先通过 Fairwind Program 面向可信网络防御者开放,价格为每百万输入 token $2、输出 token $10,缓存输入 token 为输入价的 5%。
推荐理由:官方公告给出定价、输出 token 上限和多个基准分数,读者可以据此评估它在编码与安全防御场景的落点。
Gemini 的 PostTrainBench 得分翻了一倍多:21.99%(3.1 Pro)→ 45.3%(4)🔥🚀
Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:
我让我的 dot 给我画张像。它先发来一张卡通风格的,我让它再努力点,去网上找我最近的照片。它画得好多了。 下面的视频是我在点看我的 dot 的电脑。
We’re excited to announce that @ultravox_dot_ai is now part of Inworld. Ultravox is the platform developers use to build real-time voice agents. We've worked with the team for a while through our TTS partnership, and today members of the team that built it are joining Inworld to keep developing it.
Google 发布 Gemini 4 Argon,Sundar Pichai 称其在复杂工作流、网络防御和软件工程上表现前沿。
Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:
MASSIVE reveal from Google. Its new flagship, Gemini 4 Argon, outscores GPT-6 Astra and Claude Opus 5.5 on most benchmarks. - beats GPT-6 Astra and Claude Opus 5.5 on some super important industry benchmarks. - its widest lead in legal work, 19.6% on Harvey's Legal Agent Benchmark against 6.7% for Anthropic's Claude Fable 5.1. - output limit jumps from 64K to 1M tokens, an industry-leading ceiling, - Only 3 groups have it today. the first is Google's own staff, vetted cyber defenders such as government agencies and security companies and trusted testers giving Google feedback. - Inside Google, Argon agents freed over 300 TiB of data-center memory, with 500 TiB to 1 PiB of total savings estimated, and made a Rust port of the libgav1 video decoder 2.7x faster by replacing 32K lines of SIMD code.
推荐理由:Artificial Analysis 实测显示 Gemini 4 Argon 智能指数追平 GPT-6 Astra 而折扣下成本仅六成,读者可据此比较各家模型性价比。
Google DeepMind 发布新前沿模型 Gemini 4 Argon,通过 Fairwind Program 向部分受信任测试者开放。
Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
推荐理由:榜单方公布了 Gemini 4 Argon (High) 在 Text Arena 的分项名次、1525 分和混合价格,读者可据此对比成本效率。
研究揭示在线策略蒸馏(OPD)在弱到强、同基座和强到弱师生设置下的缩放规律:早期训练中,留出准确率(gold score)随学生初始化 token 级反向 KL 散度的平方根近似线性上升。
Google 回来了??? 全面优于 Astra 和 Opus 5.5。 如果这不只是刷榜,我很想看到他们重新加入竞赛。
Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:
Robinhood 董事长兼 CEO Vlad Tenev 在 Houston 峰会上发布公司新的 AI 智能体应用,并称其安全可靠。他表示要让个人交易者获得对冲基金级别的工具,包括跨资产类别 24/7 交易、用户睡眠时仍可运行的自主“agent loops”,以及卫星影像和区块链分析等专用数据源。
Bloomberg 报道,Google 在准备发布 Gemini 4 时面临内部质疑,员工实际使用中发现该模型在编码等关键任务上表现不佳。据知情人士称,尽管 Gemini 4 在行业基准测试中成绩良好,但实际投入使用时效果不及预期,部分编码任务难以处理。
美光(Micron)给出的本财季(截至 11 月)营收指引约为 615 亿美元,高于分析师平均预期的 568 亿美元;剔除部分项目后每股利润预计约 38.15 美元,超过 36.02 美元的预期。公司称 AI 建设带来前所未有的需求,推动需求超过供给,但标题同时提到加薪压缩了利润率。
美光科技发布 2026 财年年报,营业总收入 1331.88 亿美元,同比增长 256.33%,归母净利润 849.69 亿美元,同比增长 895.07%,毛利率 80.7%。第四财季营收 542.29 亿美元,环比增长 30.81%;公司预计 2027 财年第一季度营收 600 亿至 630 亿美元,并已量产 512GB DDR5 RDIMM 内存模块,速率可达 9200 MT/s。